klip.news

What India publishes,
measured.

Klipnews is a research archive of India's digital news. This is a portrait of the corpus itself — how much is published, when, how long it runs, and how much of it is the same story reprinted.

Every figure below is measured from articles actually collected. No outlet is named anywhere on this page, deliberately.

265outlets tracked
10languages collected
94,617articles held
99outlets with articles

Collecting since 25 August 2026 — 6 days. The counts above are not stages of one queue: 265 outlets are tracked, 118 are actively crawled, and 99 have filed anything we hold. The rest are unreachable, declined by their own robots.txt, or not yet in the crawl — not merely pending. 22 outlets are actively crawled and have produced nothing, which is a gap rather than a plan. The language count is likewise what has been collected, not what those outlets publish in.

When India files its copy

Publication time, Indian Standard Time. Newsrooms have a rhythm, and it is visible in the data: the desk fills through the working day and empties overnight.

00: 4,28001: 2,29102: 91303: 94804: 1,39405: 1,25506: 2,32607: 3,22108: 4,90609: 4,57210: 4,82011: 5,15712: 5,63913: 5,28014: 5,20715: 5,59016: 5,65317: 5,27018: 4,76419: 4,60220: 4,19621: 4,04422: 3,93323: 4,356
0003060912151821
Busiest hour: 16:00 IST, with 5,653 of 94,617 articles. A large spike at 00:00 would mean publication times are being fabricated from a date-only field — a check this page runs on itself.

How long a story runs

Word counts vary enormously — a two-paragraph wire brief and a four-thousand-word investigation are both "an article".

154 words
shortest tenth
407
typical
824
longest tenth
The middle mark is the typical article. The band covers the middle 80% — the shortest and longest tenths fall outside it.

How a page gives up its text

Most publishers embed their article as structured data for search engines. Where they do, reading it is exact. Where they do not, the text has to be recovered from the page — and that is where errors creep in.

  • Structured data, with the text 80.1%
  • Structured data, text from the page 8.2%
  • Social-preview tags 5.6%
  • Read like a human would 6.1%
The first two bands are structured data the publisher provides on purpose. The last is a best-effort read of the page, and the one worth watching: if it grows for a publisher, something about their site changed.

The same story, many times

Wire agencies file once and dozens of outlets run it. Counting those as separate stories would overstate how widely anything was covered.

3.4% verbatim reprints
Identical text, published under more than one masthead. 2,307 stories account for 5,550 copies.
6.3 image links per article
Links, not pictures. The address of every image is recorded when the page is read, because those links expire long before the story does — but no image files are downloaded, so the archive holds none of the pictures themselves.

Languages in the corpus

10 languages have been collected so far, and the archive is overwhelmingly English while collection runs English first. The tracked outlets publish in more than this; those are languages we intend to reach, not languages we hold.

English92,516
Hindi1,764
Marathi269
Malayalam24
Telugu20
Tamil15
Kannada4
Urdu2
Punjabi2
Odia1

How we read the web

Politely, and checkably. Every request is logged, so what follows is measured from the log rather than quoted from the configuration — including where the two disagree.

2requests at once, per site
2sminimum gap between them
31busiest minute, any one site
3.5%answered "unchanged"
164,671requests logged
A 2-second gap is exactly 30 requests in a clock minute, so a correctly spaced run that straddles a minute boundary is counted as 31 — which is why the busiest-minute figure sits a little above the floor rather than on it, and why we publish it instead of rounding it away. Anything materially above that is a fault on our side and is investigated. We ask each site whether a page has changed before downloading it again. When the answer is no, nothing is transferred — which is why that matters to publishers as much as to us. robots.txt is honoured and re-checked weekly.