Apache Hadoop
MentionedPlatform
Also known as: Hadoop
Framework for distributed storage and processing of very large datasets across clusters of machines.
Apache Hadoop, explained
Written by EduVerseWhat it is
Apache Hadoop is an open-source framework for storing and processing data that’s too large for one machine. HDFS spreads files in blocks across many servers, with copies on several nodes. YARN shares out the cluster’s resources, and MapReduce runs computations close to where the data lives.
Why teams use it
Hadoop made it possible to process huge datasets on ordinary hardware instead of one very expensive server. These days Spark often does the processing, and many companies moved to cloud storage and data warehouses. Still, plenty of larger companies run Hadoop clusters, so there’s a fair chance you’ll meet one.
An example from work
On a data team, you need to check yesterday’s event files. You run hdfs dfs -ls /data/events/ to list them and hdfs dfs -du -h to see their size, then point a Spark job at the hdfs:// path to count events per customer.
Our own explanation, not a quote from the book.
Where it fits
Big-data storage and batch processing
Coverage in the book
Mentioned as part of the wider landscape, without in-depth coverage.