[cc] June 2026 Crawl Archive and Corresponding Web Graph are now available

84 views
Skip to first unread message

Sebastian Nagel

unread,
Jul 28, 2026, 12:33:38 PMJul 28
to Common Crawl
Hi everyone,

The July 2026 Crawl Archives and corresponding Web Graphs are now available.

The July 2026 crawl comes with several improvements and changes.
Please visit the release announcement [1] on our website.
For the Web Graphs please see [2].

Best,
Sebastian

[1] https://commoncrawl.org/blog/july-2026-crawl-archive-now-available
[2]
https://commoncrawl.org/blog/host--and-domain-level-web-graphs-may-june-and-july-2026

Julia Katash

unread,
Aug 5, 2026, 4:15:54 AMAug 5
to Common Crawl
Hello, everyone !  Useful information for all PDF users) Dual Lab company presents analysis of 20.6 Million PDF Documents from the June 2026 Common Crawl Dataset.

🚀 PDF trends 2026Q2 by Dual Lab company!

Analysis of 20.6 Million PDF Documents from the June 2026 Common Crawl Dataset!
The analysis provides a large-scale view of how PDF technology is used across the public web and establishes a foundation for future reports on PDF in general with focus on Tagged PDF, PDF/UA adoption, and accessibility trends.

Analysis https://groups.google.com/g/duallab/c/hIbFNLaqG_Y/m/UkThzNaeAAAJ


Julia Katash | PR Manager @ Dual Lab | Avenue des Villas 56/1, 1340 Ottignies-Louvain-la-Neuve, Belgium

E. julia.katash@duallab.com | http://www.duallab.com/


вторник, 28 июля 2026 г. в 19:33:38 UTC+3, Sebastian Nagel:

Sebastian Nagel

unread,
Aug 5, 2026, 4:27:39 AMAug 5
to common...@googlegroups.com
Hi Julia,

thanks for sharing the report!


One minor comment:

The report states:

> Because Common Crawl stores only the first 1 MB of each PDF,
documents > exceeding this size were downloaded directly from their
original URLs > to enable complete analysis.

The 1 MiB limit has been raised to 5 MiB in March 2025, see
https://commoncrawl.org/blog/march-2025-crawl-archive-now-available
https://commoncrawl.org/errata/content-is-truncated


I'll look deeper into the report!


Thanks again and best,
Sebastian


On 8/5/26 09:20, Julia Katash wrote:
> Hello, everyone !  Useful information for all PDF users) Dual Lab
> company presents analysis of 20.6 Million PDF Documents from the June
> 2026 Common Crawl Dataset.
>
> *🚀 PDF trends 2026Q2 by Dual Lab company!*
>
> Analysis of 20.6 Million PDF Documents from the June 2026 Common Crawl
> Dataset!
> The analysis provides a large-scale view of how PDF technology is used
> across the public web and establishes a foundation for future reports on
> PDF in general with focus on Tagged PDF, PDF/UA adoption, and
> accessibility trends.
>
> *Analysis * https://groups.google.com/g/duallab/c/hIbFNLaqG_Y/m/
> UkThzNaeAAAJ <https://groups.google.com/g/duallab/c/hIbFNLaqG_Y/m/
> UkThzNaeAAAJ>
>
> https://pdf4wcag.com/blog-news/PDF-trends-2026Q2-by-dual-lab-company
> <https://pdf4wcag.com/blog-news/PDF-trends-2026Q2-by-dual-lab-company>
>
> Julia Katash | PR Manager @ Dual Lab | Avenue des Villas 56/1, 1340
> Ottignies-Louvain-la-Neuve, Belgium
>
> E. _julia....@duallab.com_ <mailto:boris....@duallab.com> |
> _http://www.duallab.com/_ <http://www.duallab.com/>
>
>
> вторник, 28 июля 2026 г. в 19:33:38 UTC+3, Sebastian Nagel:
>
> Hi everyone,
>
> The July 2026 Crawl Archives and corresponding Web Graphs are now
> available.
>
> The July 2026 crawl comes with several improvements and changes.
> Please visit the release announcement [1] on our website.
> For the Web Graphs please see [2].
>
> Best,
> Sebastian
>
> [1] https://commoncrawl.org/blog/july-2026-crawl-archive-now-
> available <https://commoncrawl.org/blog/july-2026-crawl-archive-now-
> available>
> [2]
> https://commoncrawl.org/blog/host--and-domain-level-web-graphs-may-
> june-and-july-2026 <https://commoncrawl.org/blog/host--and-domain-
> level-web-graphs-may-june-and-july-2026>
>
> --
> You received this message because you are subscribed to the Google
> Groups "Common Crawl" group.
> To unsubscribe from this group and stop receiving emails from it, send
> an email to common-crawl...@googlegroups.com <mailto:common-
> crawl+un...@googlegroups.com>.
> To view this discussion visit https://groups.google.com/d/msgid/common-
> crawl/3aab4767-5f74-4160-b6ed-7d7eeb9a12ebn%40googlegroups.com <https://
> groups.google.com/d/msgid/common-crawl/3aab4767-5f74-4160-
> b6ed-7d7eeb9a12ebn%40googlegroups.com?utm_medium=email&utm_source=footer>.

Luca Foppiano

unread,
Aug 19, 2026, 6:04:08 AMAug 19
to Common Crawl
Hi,
  thank you for the report.

I was wondering whether you could share a) how did you process the corpus and b) the full corpus, including data and information of the re-fetched PDF documents would be made available for further study, in the same way that the SafeDocs was published (https://digitalcorpora.org/corpora/file-corpora/cc-main-2021-31-pdf-untruncated/). 

I've recently wrote a short study on SafeDocs: https://arxiv.org/pdf/2608.16390 which, among other things, recommends that number of documents are followed by number of tokens. One reason I've discovered, in the SafeDocs dataset, a short percentage of large PDF documents hold most of the tokens in the corpus.

Regards
Luca

Boris Doubrov

unread,
Aug 19, 2026, 2:36:41 PMAug 19
to common...@googlegroups.com

Hi Luca,

 

To answer your questions:

 

  1. Re-fetch truncated files

-> Apply custom Java CLI app (based on the veraPDF parser) to extract various data from PDF in JSON

-> Aggregate all JSONs into a single Parquet data file

 

              All further analysis is done in Google Colab (High-RAM runtime) based on the dataset in this single Parquet file.  

 

  1. Currently full corpus with re-fetched files is kept on a private AWS Glacier Deep Archive. But we were thinking about applying for the AWS Open Data Sponsorship to make it public. If Common Crawl has any alternative suggestions, we’d be happy to follow them.

 

Thanks a lot for the reference to your arXiv paper. We don’t (yet) tokenize the PDFs. But the file size distribution shows a similar pattern: removing only 0.1% of largest PDFs decreases the arithmetic mean of PDF file size by almost 10%. Removing 1% changes it by almost 30%. We’ll publish more details on these stats tomorrow in our blog post.

 

Best regards,

Boris

 

---------------------------------------------------------------

Boris Doubrov | CEO @ Dual Lab | Avenue des Villas 56/1, 1340 Ottignies-Louvain-la-Neuve, Belgium

F. +32 2 512 50 12 | P. +32 484 295 481 | E. boris....@duallab.com | http://www.duallab.com/

--

You received this message because you are subscribed to the Google Groups "Common Crawl" group.

To unsubscribe from this group and stop receiving emails from it, send an email to common-crawl...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/common-crawl/4b54f9a1-9f0a-46d9-b63b-439b78a34a62n%40googlegroups.com.

Greg Lindahl

unread,
Aug 20, 2026, 12:30:24 PMAug 20
to common...@googlegroups.com
Boris,

We have a "contrib" section in our free bucket, and would love to host this dataset.

-- greg


Reply all
Reply to author
Forward
0 new messages