Klipnews is a research archive of India's digital news. This is a portrait of
the corpus itself — how much is published, when, how long it runs, and how much of it is the
same story reprinted.
Every figure below is measured from articles actually collected. No outlet is
named anywhere on this page, deliberately.
265outlets tracked
10languages collected
94,617articles held
99outlets with articles
Collecting since 25 August 2026 — 6 days. The counts above are not stages of one
queue: 265 outlets are tracked, 118 are actively
crawled, and 99 have filed anything we hold. The rest are unreachable,
declined by their own robots.txt, or not yet in the crawl — not merely pending. 22
outlets are actively crawled and have produced nothing, which is a gap rather than a plan. The
language count is likewise what has been collected, not what those outlets publish in.
When India files its copy
Publication time, Indian Standard Time. Newsrooms have a rhythm, and it is
visible in the data: the desk fills through the working day and empties overnight.
0003060912151821
Busiest hour: 16:00 IST, with 5,653 of
94,617 articles.
A large spike at 00:00 would mean publication times are being fabricated from a date-only field — a check this page runs on itself.
How long a story runs
Word counts vary enormously — a two-paragraph wire brief and a
four-thousand-word investigation are both "an article".
154 words shortest tenth407 typical824 longest tenth
The middle mark is the typical article. The band covers the middle
80% — the shortest and longest tenths fall outside it.
How a page gives up its text
Most publishers embed their article as structured data for search engines.
Where they do, reading it is exact. Where they do not, the text has to be recovered from
the page — and that is where errors creep in.
Structured data, with the text 80.1%
Structured data, text from the page 8.2%
Social-preview tags 5.6%
Read like a human would 6.1%
The first two bands are structured data the publisher provides on
purpose. The last is a best-effort read of the page, and the one worth watching: if it
grows for a publisher, something about their site changed.
The same story, many times
Wire agencies file once and dozens of outlets run it. Counting those as
separate stories would overstate how widely anything was covered.
3.4%verbatim reprints
Identical text, published under more than one masthead.
2,307 stories account for 5,550
copies.
6.3image links per article
Links, not pictures. The address of every image is recorded when the
page is read, because those links expire long before the story does — but
no image files are downloaded,
so the archive holds none of the pictures themselves.
Languages in the corpus
10 languages have been collected so far, and the archive is
overwhelmingly English while collection runs English first. The tracked outlets publish in
more than this; those are languages we intend to reach, not languages we hold.
English
92,516
Hindi
1,764
Marathi
269
Malayalam
24
Telugu
20
Tamil
15
Kannada
4
Urdu
2
Punjabi
2
Odia
1
How we read the web
Politely, and checkably. Every request is logged, so what follows is measured
from the log rather than quoted from the configuration — including where the two disagree.
2requests at once, per site
2sminimum gap between them
31busiest minute, any one site
3.5%answered "unchanged"
164,671requests logged
A 2-second gap is exactly 30 requests in a clock minute, so a correctly
spaced run that straddles a minute boundary is counted as 31 — which is why the busiest-minute
figure sits a little above the floor rather than on it, and why we publish it instead of
rounding it away. Anything materially above that is a fault on our side and is investigated.
We ask each site whether a page has changed before downloading it again.
When the answer is no, nothing is transferred — which is why that matters to publishers as
much as to us. robots.txt is honoured and
re-checked weekly.