CC-News Sources Selection Question

17 views
Skip to first unread message

Giovanni MAGGI

unread,
Aug 4, 2026, 11:10:52 AM (4 days ago) Aug 4
to Common Crawl
Hi everyone, 

I have recently started to use CC-News as a main data source for my PhD project. It is a fantastic source, thank you! 

Still, I was wondering where I could find some more information on how seed sources are selected and more generally about which sites are crawled and which are not.  Is there any white paper / blog post explaining the process that goes into the collection of CC-News data? 

Thank you in advance. 

Giovanni

Sebastian Nagel

unread,
Aug 4, 2026, 1:59:47 PM (4 days ago) Aug 4
to common...@googlegroups.com
Hi Giovanni,

> It is a fantastic source, thank you!

Thank you! - It's very nice to hear that the dataset appears to be
useful. If you have any results to share, please let us know!


> Is there any white paper / blog post

Unfortunately, all information is spread over discussions in this
group and issues / PRs on GitHub. But let me try to answer your question
quickly...


Apart from few manually collected seeds, the bulk of the news crawler
seeds (RSS feeds or news sitemaps) is from DMOZ [1,2].


Some metrics about the coverage in 2021 is shared in [3].
Important to note that the picture has changed significantly
since 2021. From the 12,000 domains crawled five years ago,
about 50% are now lost - which may have multiple reasons:
1. the crawler lost the RSS feed
2. the news site disappeared
3. CCBot is now disallowed per robots.txt

Very likely, reason 3 is the dominant one. But I have no metrics at
hand to which extend. There is a paper by French researchers [4] which
shows that the news genre is the one where CCBot is disallowed the most.


We plan to upgrade the news crawler during the next months, starting
by moving it to fresh hardware during the next days. Several seed
additions are planned [5,6]. We hope that the news dataset sees some
increase in size and a more global coverage, that is news in more
languages and from more countries and geographic regions.


Best,
Sebastian

[1] https://groups.google.com/g/common-crawl/c/9JfWaFW_EQ8/m/cryNU5gnBAAJ
[2] https://github.com/commoncrawl/news-crawl/issues/8
[3] https://groups.google.com/g/common-crawl/c/SkGNdov1Mh4/m/G-NF8cxHDwAJ
[4] https://hal.science/hal-05302425/
[5] https://github.com/commoncrawl/news-crawl/issues/50
[6] https://github.com/commoncrawl/news-crawl/issues/53
Reply all
Reply to author
Forward
0 new messages