Training new font in tesseract 3.04

185 views
Skip to first unread message

Sergei Turko

unread,
Jan 23, 2017, 3:07:44 AM1/23/17
to tesseract-ocr
Hi. I'm trying to train new belarussian font in tesseract 3.04. I edited box files, made some traineddata files with dictionaries. Now I have traineddata, which do some correct recognition results, but not 80% quality at least. How can I improve recognition quality? I have multi tiff with 29 pages. Now I use this traineddata like  "-l bel+bel1"

ShreeDevi Kumar

unread,
Jan 23, 2017, 9:13:14 AM1/23/17
to tesser...@googlegroups.com
Have you tried with the 4.0 alpha version of software and new traineddata? What recognition rate do you get with that?

ShreeDevi
____________________________________________________________
भजन - कीर्तन - आरती @ http://bhajans.ramparivar.com

On Mon, Jan 23, 2017 at 1:37 PM, Sergei Turko <sting...@gmail.com> wrote:
Hi. I'm trying to train new belarussian font in tesseract 3.04. I edited box files, made some traineddata files with dictionaries. Now I have traineddata, which do some correct recognition results, but not 80% quality at least. How can I improve recognition quality? I have multi tiff with 29 pages. Now I use this traineddata like  "-l bel+bel1"

--
You received this message because you are subscribed to the Google Groups "tesseract-ocr" group.
To unsubscribe from this group and stop receiving emails from it, send an email to tesseract-ocr+unsubscribe@googlegroups.com.
To post to this group, send email to tesser...@googlegroups.com.
Visit this group at https://groups.google.com/group/tesseract-ocr.
To view this discussion on the web visit https://groups.google.com/d/msgid/tesseract-ocr/bf1d3a35-079d-4a79-99b3-9473999e2e9b%40googlegroups.com.
For more options, visit https://groups.google.com/d/optout.

Sergei Turko

unread,
Jan 24, 2017, 4:11:49 AM1/24/17
to tesseract-ocr
A little bit better, but I can't use tools for training because of its alpha-build version. How can I train tessdata with tesseract 4.0 using apps like jTessBoxEditor?

понедельник, 23 января 2017 г., 17:13:14 UTC+3 пользователь shree написал:
Have you tried with the 4.0 alpha version of software and new traineddata? What recognition rate do you get with that?

ShreeDevi
____________________________________________________________
भजन - कीर्तन - आरती @ http://bhajans.ramparivar.com

On Mon, Jan 23, 2017 at 1:37 PM, Sergei Turko <sting...@gmail.com> wrote:
Hi. I'm trying to train new belarussian font in tesseract 3.04. I edited box files, made some traineddata files with dictionaries. Now I have traineddata, which do some correct recognition results, but not 80% quality at least. How can I improve recognition quality? I have multi tiff with 29 pages. Now I use this traineddata like  "-l bel+bel1"

--
You received this message because you are subscribed to the Google Groups "tesseract-ocr" group.
To unsubscribe from this group and stop receiving emails from it, send an email to tesseract-oc...@googlegroups.com.

ShreeDevi Kumar

unread,
Jan 24, 2017, 4:54:39 AM1/24/17
to tesser...@googlegroups.com
Please read comments by Ray at https://github.com/tesseract-ocr/tesseract/issues/654

These talk of the LSTM training process. Since he is retraining, it may be helpful if you can share info about the new font that you are using so that he can include it in training.

You can try the fine-tune type training with 4.0 alpha, but seeing the number of lines required for good training, it maybe best to get the font included in training by Ray.

Or, you can retrain using jtessboxeditor, with 3.05

- excuse the brevity, sent from mobile

To unsubscribe from this group and stop receiving emails from it, send an email to tesseract-ocr+unsubscribe@googlegroups.com.

To post to this group, send email to tesser...@googlegroups.com.
Visit this group at https://groups.google.com/group/tesseract-ocr.
Reply all
Reply to author
Forward
0 new messages