There are some interesting approaches using tesseract with GNU Parallel described in the answers here [1]. If linux or os/x is not an easy option, some folks use Parallel in windows via Cygwin [2].
art
---
1. http://stackoverflow.com/questions/4962978/is-tesseract-3-00-multi-threaded
From: tesser...@googlegroups.com [mailto:tesser...@googlegroups.com]
On Behalf Of James Worldprogram
Sent: Saturday, May 02, 2015 2:37 AM
To: tesser...@googlegroups.com
Subject: [tesseract-ocr] Is there any way to speed up extraction using tesseract OCR Engine, while tiff file is having 600-700 pages?
During processing of tiff files, which are having 600 - 700 pages from Tesseract OCR engine with hocr option, we monitored that files are taking around 40 - 50 minutes.
We monitored that it is so much time for processing large files.
Do we have any way to speed up the process?
Following command is using: -
<Drive>:\Tesseract-OCR>tesseract.exe "Source_Tiff_File" "Destination_File" hocr
--
You received this message because you are subscribed to the Google Groups "tesseract-ocr" group.
To unsubscribe from this group and stop receiving emails from it, send an email to
tesseract-oc...@googlegroups.com.
To post to this group, send email to
tesser...@googlegroups.com.
Visit this group at http://groups.google.com/group/tesseract-ocr.
To view this discussion on the web visit
https://groups.google.com/d/msgid/tesseract-ocr/dde0beb4-1638-4b6b-b6d9-355ae82d769d%40googlegroups.com.
For more options, visit https://groups.google.com/d/optout.