<em inglês – ajuda pra TI Internacional> Re: Open-source server-side OCR software? (urgent)

6 views
Skip to first unread message

Daniela B. Silva

unread,
Jan 21, 2011, 11:32:52 AM1/21/11
to Matej Kurian, thac...@googlegroups.com
>> Hey Matej,

Great to hear from you.

I am forwarding your message to our "Transparency Hacker" community – let's see if anyone can help ;)

PS: Can you send us some links of those websites you presented at the Transparency Camp in Poland?

Dani

>> Oi Pessoal!

Encaminho abaixo uma mensagem do Matej Kurian, da Transparência Internacional Eslováquia, pedindo ajuda sobre alguns contratos públicos com os quais eles querem trabalhar – mas estão em PDF (boo). Alguém pode ajudar aí in english?

Aliás, esse capítulo da TI tem vários projetos bacanas – o que é ótimo, porque a Transparência Internacional não é costumeiramente muito próxima das tecnologias, e parece que isso está mudando. Pedi pra ele mandar links pra gente :)

Desculpa pela mensagem bilingue.

Beijo,

Dani


2011/1/20 Matej Kurian <kur...@transparency.sk>
Hi Daniela,

hope is going well for you.

I have an urgent question we need to deal with. One of our sponsors asked us for a project proposal on a short notice (Friday), and we are thinking of creating a searchable database of public contracts. We do not know whether it is IT-feasible and how much would it cost... Our idea main idea is to aggregate the contracts and crowdsource the control - let people look at the contracts they find interesting and check them...and report when something goes wrong.

We should not have problems getting copies of contracts (as public bodies are bound to publish them) but they are publishing them as scanned picture pdfs. We were thinking that we could OCR the documents ourselves, provided there is a budget solution (<3000 USD) available for Czech/Slovak language, ideally an open-source one. Would you know of anything? Could you refer me to someone who might know?

And finally - do you think it would be worthwhile just to download the contracts and let search engines (i.e. google) do the OCR part?

Thanks a lot,
Matej

--
Matej Kurian
programovy koordinator
Transparency International Slovensko
kur...@transparency.sk
(+421.02).5341.72.07

www.transparency.sk


Pedro Markun

unread,
Jan 21, 2011, 11:37:05 AM1/21/11
to thac...@googlegroups.com, Matej Kurian
Hi Matej,

there at least two things you could try!

One is the tesseract project. Wich is an opensource OCR framework...

The other one is to get in touch with the guys from DocumentCloud, wich has won the Knight News Challenge last year and is develloping a plataform to do exaclty that (it's based on tesseract as well)


[]'s
Pedro Markun

--
Você está recebendo esta mensagem porque se inscreveu no grupo "Transparência Hacker" dos Grupos do Google.
Para postar neste grupo, envie um e-mail para thac...@googlegroups.com.
Para cancelar a inscrição nesse grupo, envie um e-mail para thackday+u...@googlegroups.com.
Para obter mais opções, visite esse grupo em http://groups.google.com/group/thackday?hl=pt-BR.

Pedro Belasco

unread,
Jan 21, 2011, 12:13:06 PM1/21/11
to thac...@googlegroups.com, Matej Kurian
Hi Matej, 

Good to hear that other people are harassed by the common practice of publishing sensible information on scanned PDFs.

Actually I tryed to make something closely related to the arrange you proposed. In the occasion, we're trying to scan legal documents of our lawmakers.

The OCR scanning made by tesseract is pretty good, but have serious problems identifying text columns and spiting the text sequentially. It's a very memory intensive computational procedure, which needs a lot of computational resources to run on large bases.
The best results you can achieve, unfortunately are using the OCR libs present in the adobe's acrobat pro.

There's another technical caveat involved in this process, that's the accuracy of the results. They are somehow unpredictable, and necessarily  need to be reviewed by human hands.

The document cloud is a pretty nice initiative, but i really don't know if the libs they're working on can solve our problem.
For me, this issue it's a kind of rosetta's stone of our crusade.

Regards.

Pedro Belasco

Ricardo Poppi

unread,
Jan 24, 2011, 8:50:49 AM1/24/11
to thac...@googlegroups.com, Matej Kurian
Compilei o tesseract-3.00 num lenny e dei uma brincada. A versão 2.0 realmente melava 2 colunas, mas a 3.00 me impressionou.

Estou anexando o documento em TIFF que utilizei e o output txt obtido.

Usei esse tutorial para instalar essa versão.

Abcs!

2011/1/21 Pedro Belasco <pbel...@gmail.com>
outputtext.txt
1cmfile.tiff

Daniela B. Silva

unread,
Jan 24, 2011, 12:38:37 PM1/24/11
to thac...@googlegroups.com
Alguns links da TI Eslováquia.

Bjs

---------- Forwarded message ----------
From: Matej Kurian <kur...@transparency.sk>
Date: Sat, Jan 22, 2011 at 12:00 PM
Subject: Re: <em inglês – ajuda pra TI Internacional> Re: Open-source server-side OCR software? (urgent)
To: "Daniela B. Silva" <daniel...@gmail.com>


Hi Dani,

thanks a lot for your help. I guess we might need to scale down the project a bit and see whether we can improve it later on - but hey, that's life :)

The sites are -

http://vestnik.transparency.sk/ for public procurement (Slovak only at the moment, sorry :/)
http://samosprava.transparency.sk/en/ for the local municipalities survey

There workings of the procurement site are described here -  http://www.scribd.com/doc/35713180/Slovak-Public-Procurement-Announcements-Extraction-Transformation-and-Loading

How are you? Any exciting project going on?

Cheers,
Matej
Reply all
Reply to author
Forward
0 new messages