TL;DR

  • Hacker News (HN) is a popular tech news aggregator, and I was curious about its ranking dynamics and how virality works on HN.
  • I built an event-driven data pipeline on AWS to collect every event that happens on Hacker News.
  • I ran causal inference using a Bayesian approach and found that 1) there is an upvote-rate boost between page 1 and page 2, and 2) submitting in the morning gives you a higher chance of reaching page 1. Note that these results hold only under certain assumptions.

You can take a look at the details of the data pipeline and the analysis! You can also check out the detailed analysis report here.

What is Hacker News?

Hacker News is a social news aggregator focused on technology, startups, and computing. It was launched in February 2007 by Paul Graham as a side project.1 Users submit links to news, blog posts, papers, repositories, and other writing that fits HN’s focus on intellectual curiosity.2 Other users can upvote posts they find interesting, increasing their visibility, and leave comments in threaded discussions.

The community has attracted a loyal following for years. According to Similarweb, news.ycombinator.com had an estimated 13.7 million visits in April 2026.3 In the dataset I collected during April 2026, I observed 30,088 story submissions, or about 1,003 submissions per day. Topics change over time. As of 2026, AI-related posts make up a visible share of the conversation, but foundational topics such as compilers, operating systems, networking, and programming languages still show up frequently.

Why bother?

First, I am a heavy HN user. I read Hacker News whenever I get bored. Second, going viral on HN, which I define here as reaching the top 10 of the main page, is known to send substantial traffic to the linked page. There are plenty of stories like this: one author reported more than 45,000 requests and more than 29,000 unique IP addresses in one day after HN exposure, while another Reddit post reported 7.5k first-day visits from a Show HN launch.45 Many people treat seeing their side project on the HN front page as an honor, and the value of HN virality is widely recognized.6 As an applied ML enthusiast, it was natural for me to wonder how a product I use every day creates virality.

As I looked into it, I learned that HN is maintained through the work of moderators, especially the well-known dang. Moderation is not just about blocking guideline violations in posts and comments. It also involves substantial curation effort, including giving high-quality buried submissions a second chance through mechanisms such as the second-chance pool.67

In other words, HN is shaped by the dynamics among three parties: users who try to submit good posts, users who evaluate posts by voting, and moderators who intervene to maintain the overall quality of the ecosystem. This makes HN an interesting place to study virality: the front page is shared, ranking is visible, and votes and comments leave enough traces to study how visibility and engagement reinforce each other.

How

Long story short, I built my own data pipeline to capture snapshots of each Hacker News post and ranking state. This was possible because the official Hacker News API exposes the current state of items and the current Top Stories list.8 I wrote more about the pipeline design in Hacker News Data Pipeline.

After building the pipeline, I could analyze the life cycle of each post. From the moment a story was submitted, I had records of how many upvotes it received, when its rank moved up or down, when it crossed page boundaries, and how long it stayed visible.

For the causal analysis, I turned those records into a few explicit hypotheses and a causal DAG. Under that DAG, I used Bayesian inference to estimate whether page 1 exposure increased later upvote rate, and whether submission timing affected the chance of reaching the Top 30. The details are in Causal Analysis of Virality on Hacker News.

Results

The data pipeline worked as expected and collected data for almost six months without a major operational issue. Designing and operating the pipeline was a useful learning experience by itself. I wrote the operational details separately in Hacker News Data Pipeline.

The analysis also confirmed the main intuition. Reaching the Top 30, which means getting onto the first page, appears to change the amount of additional upvotes a post receives. The timing result pointed in the same direction as my intuition. Submitting in the morning looked more favorable for entering the Top 30.

Those results should be read with care. They depend on my causal DAG being a reasonable approximation of reality. They also assume no important moderator intervention or unobserved confounder is driving the result. Those are somewhat naive assumptions for a system like HN, so I would not treat the estimates as a perfect description of the real world. The details and caveats are in the causal analysis write-up.

Next Steps

With the data pipeline now in place, there is much more analysis to run. First, the causal analysis could be improved. For example, I could extract and use semantic features from titles and content. The causal DAG could also be refined; the current analysis relies on a naive assumption without feedback loops.

It is also possible to personalize Hacker News. The current pipeline does not include individual user engagement events, but it can still model average community-level preferences. That community-level model could serve as a seed model and be fine-tuned with a small amount of individual user data. For example, I could collect my own local read-history data and use it to fine-tune a personalized model for myself. Based on that, one possible exploration is to surface posts that a user would likely enjoy but that failed to go viral.

Another direction is community moderation support. Today, HN is maintained effectively through human moderators’ work. It would be interesting to see whether a system can learn moderation patterns from data and help human moderators more effectively.

References