Hi Matej,
Good to hear that other people are harassed by the common practice of publishing sensible information on scanned PDFs.
Actually I tryed to make something closely related to the arrange you proposed. In the occasion, we're trying to scan legal documents of our lawmakers.
The OCR scanning made by tesseract is pretty good, but have serious problems identifying text columns and spiting the text sequentially. It's a very memory intensive computational procedure, which needs a lot of computational resources to run on large bases.
The best results you can achieve, unfortunately are using the OCR libs present in the adobe's acrobat pro.
There's another technical caveat involved in this process, that's the accuracy of the results. They are somehow unpredictable, and necessarily need to be reviewed by human hands.
The document cloud is a pretty nice initiative, but i really don't know if the libs they're working on can solve our problem.
For me, this issue it's a kind of rosetta's stone of our crusade.
Regards.
Pedro Belasco