Skip to content
EduVerse
Apache Hadoop

Apache Hadoop

Mentioned

Platform

Also known as: Hadoop

Framework for distributed storage and processing of very large datasets across clusters of machines.

Apache Hadoop, explained

Written by EduVerse

What it is

Apache Hadoop is an open-source framework for storing and processing data that’s too large for one machine. HDFS spreads files in blocks across many servers, with copies on several nodes. YARN shares out the cluster’s resources, and MapReduce runs computations close to where the data lives.

Why teams use it

Hadoop made it possible to process huge datasets on ordinary hardware instead of one very expensive server. These days Spark often does the processing, and many companies moved to cloud storage and data warehouses. Still, plenty of larger companies run Hadoop clusters, so there’s a fair chance you’ll meet one.

An example from work

On a data team, you need to check yesterday’s event files. You run hdfs dfs -ls /data/events/ to list them and hdfs dfs -du -h to see their size, then point a Spark job at the hdfs:// path to count events per customer.

Our own explanation, not a quote from the book.

Where it fits

Big-data storage and batch processing

Coverage in the book

Mentioned

Mentioned as part of the wider landscape, without in-depth coverage.

Appears in