poor tesseract training

163 views
Skip to first unread message

Lada Tylich

unread,
Oct 25, 2017, 4:18:52 PM10/25/17
to tesseract-ocr
Hi to All!

I try to train tesseract to "computer-like" and "digital-like" fonts (rugsnatcher-demo and levi-adobe-dia in particular). I kept the training recommendations for image preprocessing, changing box file using qt-box-editor, spaces between characters. But despite these I got fails when I put the trainign command (e.g. tesseract eng.rugsnatcher_demo.exp0.tif eng.rugsnatcher_demo.exp0 box.train). In some cases I got through but finally at combine_tessdata procedure (Error: traineddata file must contain at least (a unicharset fileand inttemp) OR an lstm file.).

Editing box files seems totally useless to me since changing 1 "failing" character (that is correct btw) makes even more fails during training.

I have already tried another font sf_digital_readouts. Considering this, training was successful only for few characters, but when I put there paragraph of text I got some Fails as well. 

I am stuck at this few weeks but I am not able to make good results for this.  I am curious what am I missing? Is it possible to teach tesseract these fonts or should I take another approach?

Thank you for any response!

ShreeDevi Kumar

unread,
Oct 25, 2017, 11:16:44 PM10/25/17
to tesser...@googlegroups.com
If you have the font files, best way to do training is to use tesseract-ocr/tesseract/training/tesstrain.sh - it will create the box/tiff pairs and the resulting .lstmf files (for 4.0 alpha training) or .tr files (for 3.0x training).

ShreeDevi
____________________________________________________________
भजन - कीर्तन - आरती @ http://bhajans.ramparivar.com

--
You received this message because you are subscribed to the Google Groups "tesseract-ocr" group.
To unsubscribe from this group and stop receiving emails from it, send an email to tesseract-ocr+unsubscribe@googlegroups.com.
To post to this group, send email to tesser...@googlegroups.com.
Visit this group at https://groups.google.com/group/tesseract-ocr.
To view this discussion on the web visit https://groups.google.com/d/msgid/tesseract-ocr/de9d3d76-115b-42f8-a80a-806588d85059%40googlegroups.com.
For more options, visit https://groups.google.com/d/optout.

Lada Tylich

unread,
Oct 26, 2017, 5:53:10 AM10/26/17
to tesseract-ocr
Thanks for answer.
I ran the command (tesstrain.sh). It was able to generate training images but it failed during generating unicharset and unichar properties files:

Invalid Unicode codepoint: 0xffffffc2
IsValidCodepoint(ch):Error:Assert failed:in file normstrngs.cpp, line 225
ERROR: /tmp/tmp.TW7Mm4rzdo/eng/eng.unicharset does not exist or is not readable

I checked that it can relate to some older commits but I got them already. I try to find where is problem.

Do you know what else can cause the issue?




 

Lada Tylich

unread,
Oct 26, 2017, 6:15:33 AM10/26/17
to tesseract-ocr
btw: I use tesseract commit 1b0379c

Lada Tylich

unread,
Oct 26, 2017, 7:45:38 AM10/26/17
to tesseract-ocr
and I tried different fonts and English language, so there should not be issues like here.

Lada Tylich

unread,
Oct 29, 2017, 2:11:20 PM10/29/17
to tesseract-ocr
So please, referring to my previous comments, do you have some ideas? I am stuck again :/
Reply all
Reply to author
Forward
0 new messages