TL;DR
- I built a data pipeline that stores public events from Hacker News.
- The first goal was data analysis, but I also wanted a reusable event backbone for future HN projects, including recommendation systems.
- Tech Stack: EC2, SQS, Lambda, DynamoDB, Kinesis Data Firehose, S3, Athena, CloudWatch, Terraform, and Docker.
The data collected by this pipeline was used in When Does a Post Go Viral on Hacker News?.
Why I Built This
Hacker News already has a public API.1 It exposes a useful view of the current HN ecosystem, including item details, user profiles, recently changed items and profiles, and the current Top Stories list.
That is enough for many applications. If I want to show the current title, score, and comments for a story, the API is fine. For the analysis I had in mind, though, the current state was not enough. I was interested in the dynamics of the system. I wanted to know what path a post took before it reached the Top 10, how its upvotes accumulated, when comments arrived, when it moved between pages, and whether author-level signals were useful later.
Those questions need snapshots over time. A final item record does not tell me how the item got there.
So I decided to build my own pipeline. Part of the motivation was practical. I needed the data for analysis. I also wanted to practice building something closer to a real data pipeline than a local scraper.
For Hacker News scale, the simplest system might have been one EC2 instance that polls the API and writes files directly to disk or S3. That probably would have been cheaper and easier. I intentionally chose a more overbuilt design because I wanted to practice the shape of a larger production data system. That meant queues, serverless workers, latest-state storage, append-only archives, CDC, monitoring, and infrastructure-as-code.
That made the project more complicated than it strictly needed to be. The extra complexity gave me a reusable base for later HN work. Once I have item snapshots, ranking events, and user snapshots in a consistent format, I can attach other pipelines on top for causal analysis, recommender experiments, author features, topic modeling, or personalized ranking.
Hacker News API Overview
The HN API is small, but it has the pieces needed to observe the public surface of the site.
The first entity is an item. In HN, stories, comments, jobs, polls, and poll options are all items. They live under this endpoint.
/v0/item/<id>.json
An item can have fields like id, type, by, time, title, url, text, score, kids, descendants, dead, and deleted. This is the main entity for the item pipeline.
The second entity is a user.
/v0/user/<id>.json
A user record contains fields like id, created, karma, about, and submitted. I used this for author context and future recommender features.
The API also exposes live-ish change surfaces.
/v0/updates.json
/v0/topstories.json
updates.json returns recently changed item IDs and profile IDs. I used it as the source for the item and user pipelines. topstories.json returns the current Top Stories ranking list, up to 500 items. I used it to build ranking snapshots and ranking events.
Item Pipeline

Figure 1. Item pipeline from HN updates to raw snapshots and derived change events.
The item pipeline is the main pipeline.
It starts with an EC2 poller. The poller reads updates.json, looks at the items array, and detects item IDs that were not in the previous poll. It sends those IDs to SQS.
The next step is a Lambda fetcher. The fetcher reads item IDs from SQS, calls /v0/item/<id>.json, computes a deterministic hash of the HN payload, and writes the latest item state to DynamoDB only if the payload changed. When it writes a changed item, it also sends the raw item snapshot to Firehose, which writes compressed JSONL files into S3.
DynamoDB Streams are enabled on the latest-items table. A second Lambda receives old and new images from the stream, compares the item field, and emits field-level change events to another Firehose stream. Those events also land in S3.
So the item pipeline has two S3 outputs.
raw item snapshots
item change events
This split was deliberate. The raw snapshot is the safer source of truth. The event stream is a derived convenience layer. If I later decide that an event definition was wrong, I can recompute it from raw snapshots.
The advantage of this design is that each piece has a simple job. The poller detects changes. SQS absorbs bursts. Lambda fetches item details. DynamoDB holds the latest state. S3 holds history. The CDC Lambda turns state changes into analysis-friendly events.
The downside is obvious. This is more moving parts than HN strictly requires. It also has a small blind spot around poller restarts because the poller keeps its previous baseline in memory. If the poller restarts, it takes a fresh baseline and only catches later diffs. I accepted this because the goal was not perfect archival of every possible API state. I wanted a reliable enough public-event history for analysis.
Top Stories Pipeline

Figure 2. Top Stories pipeline for rank snapshots and rank movement events.
The Top Stories pipeline tracks visibility.
The HN item API tells me an item’s score and metadata, but it does not tell me when the item was on page 1, when it crossed from rank 31 to rank 30, or how long it stayed near the top. For that, I needed ranking snapshots.
The top stories poller reads topstories.json on a fixed interval. For each snapshot, it writes the ranked list to Firehose and S3. It also compares the current ranking with the previous ranking and emits rank events.
APPEAR
DISAPPEAR
RANK_UP
RANK_DOWN
The advantage is that ranking dynamics become explicit. I can ask when an item first appeared in Top Stories, whether it reached the Top 30, how long it stayed visible, and how rank movement related to later upvotes.
The tradeoff is storage volume. A full ranking snapshot writes many records even when the ranking only changes a little. In this pipeline, that tradeoff was acceptable because the records are simple, compressed S3 storage is cheap, and the analysis becomes much easier. I would rather store a slightly redundant ranking history than reconstruct visibility from partial state later.
User Pipeline

Figure 3. User pipeline for profile snapshots and historical author backfill.
The user pipeline came later and had the lowest priority.
Conceptually, it is similar to the item pipeline. A user poller reads the profiles field from updates.json and sends changed user IDs to SQS. A Lambda fetcher calls /v0/user/<id>.json, stores the latest user profile in DynamoDB, and archives raw snapshots to S3 through Firehose.
The main difference is that I did not build a CDC event pipeline for users.
User records have much less dynamic structure than items. For my use case, I mostly needed author context such as karma, account age, profile text, and submitted history. Field-level user events were not important enough to justify another stream.
The biggest operational decision was user backfill. Since the user pipeline was added after the item pipeline, I needed a backfill to recover author profiles for historical items. I used a residual backfill instead of a blind one to avoid redundant HN API calls for users already collected by the live pipeline. I started with small canaries, then 100 users, then 1,000, then 5,000, and eventually continued in 5,000-user chunks. The rate ceiling stayed at 5 users per second. The pace was conservative. It kept the live queue healthy and avoided unnecessary load on HN.
Observability and Operation
The operational goal was not to build a perfect monitoring system. I wanted enough observability to know whether the pipeline was alive, whether queues were backing up, and whether data was still reaching S3.
The pollers publish CloudWatch metrics such as heartbeat, poll latency, and number of enqueued IDs. SQS queue age is the main downstream health signal. If the poller heartbeat is present while the queue age is rising, the fetcher side is not keeping up.
I also added a small monitoring layer with CloudWatch alarms, SNS email alerts, and a weekly summary Lambda. The summary checks monitored surfaces, throughput, alarm history, queue health, and coarse S3 freshness. It can tell me that the monitored surfaces look healthy, but it cannot prove every individual record was processed exactly as intended. That was the right level for a small always-on research pipeline.
Infrastructure was deployed with Terraform where I wanted a repeatable control plane. The pollers ran as Docker containers under systemd on EC2. Lambda deployment used scripts and packaged zip artifacts. S3 data was queried with Athena when I needed quick checks or lightweight analysis. I also added small cost controls later, including batched CloudWatch metric publication, lower heartbeat frequency, and log retention.
Results and Takeaway
The archive contains item data from November 17, 2025 UTC through May 11, 2026 UTC. The Top Stories archive starts on November 27, 2025 UTC. The user archive starts on March 18, 2026 UTC.
Across that period, CloudWatch and S3 show roughly the following volume.
- 9.4M raw item snapshot records sent to Firehose
- 10.6M item change event records sent to Firehose
- 218.2M Top Stories rank snapshot records sent to Firehose
- 29.0M Top Stories rank event records sent to Firehose
- 2.7M raw user snapshot records sent to Firehose
- 26.0M item SQS messages received
- 2.8M user SQS messages received
Operationally, I would not claim literal zero errors. CloudWatch shows a small number of transient Lambda errors across millions of invocations. I did not see a sustained pipeline incident in the metrics I checked. Firehose failed-put metrics were zero, and max queue age stayed bounded during the archived period.
I think this was partly because the system was not actually very complex. HN is also a stable community site rather than a high-variance traffic source. The traffic is large enough to be interesting and small enough that a conservative AWS pipeline can handle it comfortably.