How do I remove extra unnecessary characters from OCR output

112 views
Skip to first unread message

karra...@gmail.com

unread,
Nov 25, 2017, 4:06:17 PM11/25/17
to tesseract-ocr
Hi,

I'm using OCR for my Raspberry Pi book reader project and I seem to have this issue where the OCR produces extra characters along with the text every time. I tried using some of the config parameters to see if it makes any difference which hasn't been the case so far. 

For example, here is my output for the following text from a Dr.Seuss book: (The text that was on the page was just Everyone wants

a big green kangaroo.) How can I get rid of the extra characters? Thanks


4441‘ 'muw-u»


A; Wmme

\ K r'f'.

,\

3 I ,

E v \y r

3 \g\\ A  '

'3 I» ’1,” ‘

 ,x


i, 1/ /


I

¢ 1


“.

/ ‘x

m- WWW-““a”..-


  Everyone wants

a big green kangaroo.



: is C-Bpfright Cnnwnucns.


r-wwrwmlaai.vn~is&‘-M'ikiww. _ . A

TAPAN SHARMA

unread,
Jan 17, 2018, 12:15:36 PM1/17/18
to tesseract-ocr
It may be possible due to compression artifacts in Image. You can achieve a lot of improvement using lossless compression like PNG than JPEG. 
Reply all
Reply to author
Forward
0 new messages