Mozilla Telemetry: In-depth Data Pipeline

Mozilla Telemetry: In-depth Data Pipeline (via) Detailed behind-the-scenes look at an extremely sophisticated big data telemetry processing system built using open source tools. Some of this is unsurprising (S3 for storage, Spark and Kafka for streams) but the details are fascinating. They use a custom nginx module for the ingestion endpoint and have a “tee” server written in Lua and OpenResty which lets them route some traffic to alternative backend.

Posted 12th April 2018 at 3:44 pm

Simon Willison’s Weblog

Recent articles

Monthly briefing