Introduction to Apache Flume in 30 minutes

What is Apache Flume?

Apache Flume is a distributed, reliable, and available system for efficiently collecting, aggregating & moving large data from many different sources to a centralized data store.

Flume supports a large variety of sources Including:

tail (like unix tail -f),
syslog,
log4j – allowing Java applications to write logs to HDFS via flume

Flume Nodes

Flume nodes can be arranged in arbitrary topologies.Typically there is a node running on each source machine, with tiers of aggregating nodes that the data flows through on its way to HDFS.

Topics Covered

What is Flume
Flume: Use Case
Flume: Agents
Flume: Use Case – Agents
Flume: Multiple Agents
Flume: Sources
Flume: Delivery Reliability
Flume: Hands-on

Introduction to Flume Presentation

Please feel free to leave your comments in the comment box so that we can improve the guide and serve you better. Also, Follow CloudxLab on Twitter to get updates on new blogs and videos.

If you wish to learn Hadoop and Spark technologies such as MapReduce, Hive, HBase, Sqoop, Flume, Oozie, Spark RDD, Spark Streaming, Kafka, Data frames, SparkSQL, SparkR, MLlib, GraphX and build a career in BigData and Spark domain then check out our signature course on Big Data with Apache Spark and Hadoop which comes with

Online instructor-led training by professionals having years of experience in building world-class BigData products
High-quality learning content including videos and quizzes
Automated hands-on assessments
90 days of lab access so that you can learn by doing
24×7 support and forum access to answer all your queries throughout your learning journey
Real-world projects
A certificate which you can share on LinkedIn

What is Apache Flume?

Flume supports a large variety of sources Including:

Flume Nodes

Leave a Reply Cancel reply