Julia Katash | PR Manager @ Dual Lab | Avenue des Villas 56/1, 1340 Ottignies-Louvain-la-Neuve, Belgium
E. julia.katash@duallab.com | http://www.duallab.com/
Hi Luca,
To answer your questions:
-> Apply custom Java CLI app (based on the veraPDF parser) to extract various data from PDF in JSON
-> Aggregate all JSONs into a single Parquet data file
All further analysis is done in Google Colab (High-RAM runtime) based on the dataset in this single Parquet file.
Thanks a lot for the reference to your arXiv paper. We don’t (yet) tokenize the PDFs. But the file size distribution shows a similar pattern: removing only 0.1% of largest PDFs decreases the arithmetic mean of PDF file size by almost 10%. Removing 1% changes it by almost 30%. We’ll publish more details on these stats tomorrow in our blog post.
Best regards,
Boris
---------------------------------------------------------------
Boris Doubrov | CEO @ Dual Lab | Avenue des Villas 56/1, 1340 Ottignies-Louvain-la-Neuve, Belgium
F. +32 2 512 50 12 | P. +32 484 295 481 | E. boris....@duallab.com | http://www.duallab.com/
--
You received this message because you are subscribed to the Google Groups "Common Crawl" group.
To unsubscribe from this group and stop receiving emails from it, send an email to
common-crawl...@googlegroups.com.
To view this discussion visit
https://groups.google.com/d/msgid/common-crawl/4b54f9a1-9f0a-46d9-b63b-439b78a34a62n%40googlegroups.com.