How News Clustering Works
AllTop ingests articles from a large set of news feeds. Each incoming article is converted into a numerical representation of its meaning, built from the headline and summary, and compared against the stories already forming. If the new article is a very close match to an existing story cluster, it attaches to that cluster. If it's a moderately close match and shares a key entity with the cluster, the same person, company, or event at the center of it, it attaches too, with an additional automated check confirming borderline cases. If neither condition holds, the article starts a new story.
AllTop ingests articles from a large set of news feeds. Each incoming article is converted into a numerical representation of its meaning, built from the headline and summary, and compared against the stories already forming. If the new article is a very close match to an existing story cluster, it attaches to that cluster. If it's a moderately close match and shares a key entity with the cluster, the same person, company, or event at the center of it, it attaches too, with an additional automated check confirming borderline cases. If neither condition holds, the article starts a new story.
Clusters have a shelf life. A story older than 36 hours stops accepting new articles, so a fresh wave of coverage on a developing situation forms a new story rather than piling onto a stale one. That keeps the feed reflecting what's happening now instead of endlessly extending old clusters.
What the source count tells you
Because the count measures distinct outlets, it works as a quick corroboration gauge. A story showing one source is a single outlet's claim that nobody else has matched yet. It might be a genuine scoop, or it might not hold up. A story showing fifteen sources is an event the press corps broadly agrees happened. AllTop's ranking model treats these very differently: single-source stories carry a penalty until corroborated, and high-source stories decay more slowly at the top of the page. Our explainer on how stories are ranked covers it in full.
How to use the count
Rising count on a breaking story — coverage is spreading, which is often when a linked market starts moving.
Stubbornly low count on a dramatic claim — Nobody else has matched it yet.
Very high count on a fading story — the event was major, but the news cycle has largely finished with it.
The count is also honest about wire copy. When many outlets run near-identical syndicated text, the ranking model dampens the credit that breadth would normally earn, because ten copies of one wire story represent one act of reporting, not ten. The displayed source count still reflects the distinct feeds, but the story doesn't rank as if ten newsrooms independently confirmed it.
Why group at all
Without clustering, a major event would flood the feed with twenty separate near-identical headlines, and you'd scroll through repetition instead of news. Grouping collapses the repetition into one entry, shows you how widely it's being covered, and frees the rest of the page for stories you haven't seen. Each story entry also gets a short automated summary and a content-type label, distinguishing breaking news from analysis, features, explainers, and investigations, so you know what kind of read you're clicking into.