Distributed storage · 2003 · Sanjay Ghemawat, Howard Gobioff & Shun-Tak Leung

The Google File System

Redesign storage around measured workload facts: routine component failure, huge files, streaming reads, append-heavy writes, and commodity machines.

The central move

Redesign storage around measured workload facts: routine component failure, huge files, streaming reads, append-heavy writes, and commodity machines.

Why it had to exist

Google’s data processing exceeded the assumptions of general distributed filesystems. Failures were common enough to become a normal operating condition rather than an exception.

Where it leads

Failure-aware storage → MapReduce and Bigtable → modern data infrastructure and cloud object stores.

Study the guided reading in Bits →