संस्कृत PDF copy-paste / OCR कर्तुं सरलः उपायः

53 views
Skip to first unread message

Lokesh Sharma

unread,
Jun 7, 2023, 3:06:05 PM6/7/23
to sanskrit-programmers
नमोनमः

वयम् सर्वे जानीमः यत् देवनागर्याम् लिखितः PDF लेखः copy - paste कर्तुं बह्वी समस्या।

तस्मै एकं सरलम् उपायं प्राप्तवान् अहम् अद्य।


एवं कार्यं करोति एतत् -

$> ocrmypdf -l san input.pdf output.pdf --force-ocr

तावत् एव।

सुखिनः भवन्तु।


विश्वासो वासुकिजः (Vishvas Vasuki)

unread,
Jun 7, 2023, 8:32:09 PM6/7/23
to sanskrit-p...@googlegroups.com, Shree Devi Kumar
tesseract-data-san इति प्रयुङ्क्त इति भाति 

--
You received this message because you are subscribed to the Google Groups "sanskrit-programmers" group.
To unsubscribe from this group and stop receiving emails from it, send an email to sanskrit-program...@googlegroups.com.
To view this discussion on the web visit https://groups.google.com/d/msgid/sanskrit-programmers/CA%2BksRZcyrtz_bJhAgOMncAJZ%3Dy7D%3DZkJoE%2B8hBm5xKswXHbdeA%40mail.gmail.com.


--
--
Vishvas /विश्वासः

विश्वासो वासुकिजः (Vishvas Vasuki)

unread,
Jul 20, 2026, 6:42:22 AM (10 days ago) Jul 20
to sanskrit-p...@googlegroups.com, Anik tivAri अनीकः श्रीवैष्णवः शाण्डिल्यः
https://ocr.sanchaya.net/ uses tesseract - may be a good option.

Anunad Singh

unread,
Jul 20, 2026, 11:35:07 PM (10 days ago) Jul 20
to sanskrit-p...@googlegroups.com
Sanchaya : A great project with clear vision!

Now about my first experience of its OCR. I tried it on a text containing Sanskrit, Hindi and English. Accuracy seems as good as it has been for Tesseract for these languages. Speed is also moderate. I think processing is being done on the client side. 

The menu is easy to grasp. The options are quite many !

Some improvements i would like to suggest : 
(1) The find-replace option could also have option for regular expressions. This is important because OCRed text has good chances of having systematic errors instead of random errors. 
(2) some way to tell about selecting a table and telling about the number of columns. (I did not try on pages with two or three columns of text)
(3) It left two initial words unboxed. There could be some way to force it to 'see' those texts.
(4) 'train data' has no response. Is it active in the present version?
(5) It shows 'Page 10 of 24' at the top of left hand panel. But in reality, it displays two pages (page 10 & 11) and produces text for these displayed pages if 'recognise' is clicked. Instead of 'recognise' it could be 'recognise selected'  or something like that.

-- अनुनाद 

विश्वासो वासुकिजः (Vishvas Vasuki)

unread,
Jul 21, 2026, 12:16:41 AM (10 days ago) Jul 21
to sanskrit-p...@googlegroups.com, Omshivaprakash scanner शिवप्रकाशः पुस्तकचित्रीकृत् कल्याणपुर्याम्

omshiva...@gmail.com

unread,
Jul 22, 2026, 8:33:53 PM (8 days ago) Jul 22
to विश्वासो वासुकिजः (Vishvas Vasuki), sanskrit-p...@googlegroups.com
Thank you. Reply inline. 

On Tue, Jul 21, 2026 at 9:46 AM विश्वासो वासुकिजः (Vishvas Vasuki) <vishvas...@gmail.com> wrote:
+omshivaprakash the author

On Tue, 21 Jul 2026 at 09:05, Anunad Singh <anu...@gmail.com> wrote:
Sanchaya : A great project with clear vision!

Now about my first experience of its OCR. I tried it on a text containing Sanskrit, Hindi and English. Accuracy seems as good as it has been for Tesseract for these languages. Speed is also moderate. I think processing is being done on the client side. 

The menu is easy to grasp. The options are quite many !

Some improvements i would like to suggest : 
(1) The find-replace option could also have option for regular expressions. This is important because OCRed text has good chances of having systematic errors instead of random errors. 

That's a good point to incorporate. This tool has been evolving. 

(2) some way to tell about selecting a table and telling about the number of columns. (I did not try on pages with two or three columns of text

We don't have column/table segmentation yet. 

(3) It left two initial words unboxed. There could be some way to force it to 'see' those texts.

You can use VietOCR for various options like OCRing a selected area, line etc. 
 
(4) 'train data' has no response. Is it active in the present version?

That's the output we export for our trainocr project. As I mentioned, this project has different use-cases for different users. I shall get the documentation and guide updated. 
 
(5) It shows 'Page 10 of 24' at the top of left hand panel. But in reality, it displays two pages (page 10 & 11) and produces text for these displayed pages if 'recognise' is clicked. Instead of 'recognise' it could be 'recognise selected'  or something like that.


The text area loads the entire PDF/image first. Once its loaded you choose recognize to OCR one page recognize all for OCRing all pages. Then you click on different pages on left hand side section to keep moving to next page and its output. 


--
--
With Best Regards,
Omshivaprakash.H.L | ಓಂ ಶಿವಪ್ರಕಾಶ್ ಎಚ್. ಎಲ್ | ॐ शिवप्रकाश् एच्. एल्

ServantsOfKnowledge
CEO & Founding Trustee 
Reply all
Reply to author
Forward
0 new messages