RE: [tesseract-ocr] Is there any way to speed up extraction using tesseract OCR Engine, while tiff file is having 600-700 pages?

186 views
Skip to first unread message

Art Rhyno.

unread,
May 2, 2015, 7:41:51 AM5/2/15
to tesser...@googlegroups.com

There are some interesting approaches using tesseract with GNU Parallel described in the answers here [1]. If linux or os/x is not an easy option, some folks use Parallel in windows via Cygwin [2].

 

art

---

1. http://stackoverflow.com/questions/4962978/is-tesseract-3-00-multi-threaded

2. http://blogs.msdn.com/b/hpctrekker/archive/2013/03/30/preparing-and-uploading-datasets-for-hdinsight.aspx

 

From: tesser...@googlegroups.com [mailto:tesser...@googlegroups.com] On Behalf Of James Worldprogram
Sent: Saturday, May 02, 2015 2:37 AM
To: tesser...@googlegroups.com
Subject: [tesseract-ocr] Is there any way to speed up extraction using tesseract OCR Engine, while tiff file is having 600-700 pages?

 

During processing of tiff files, which are having 600 - 700 pages from Tesseract OCR engine with hocr option, we monitored that files are taking around 40 - 50 minutes.

We monitored that it is so much time for processing large files.

Do we have any way to speed up the process?

Following command is using: -

<Drive>:\Tesseract-OCR>tesseract.exe "Source_Tiff_File" "Destination_File" hocr

--
You received this message because you are subscribed to the Google Groups "tesseract-ocr" group.
To unsubscribe from this group and stop receiving emails from it, send an email to tesseract-oc...@googlegroups.com.
To post to this group, send email to tesser...@googlegroups.com.
Visit this group at http://groups.google.com/group/tesseract-ocr.
To view this discussion on the web visit https://groups.google.com/d/msgid/tesseract-ocr/dde0beb4-1638-4b6b-b6d9-355ae82d769d%40googlegroups.com.
For more options, visit https://groups.google.com/d/optout.

Reply all
Reply to author
Forward
0 new messages