Hi Todd,
Happy to lay it out. The short version is that the "OCR" step turned out to be the least reliable part, and I ended up leaning on a vision model reading the page images rather than a classic OCR engine. Here's the process and where the problems were.
Tools
- poppler (command-line): pdfinfo for page count/metadata, pdftotext to pull any embedded text layer, pdftoppm to rasterize pages to PNG. On a Mac that's brew install poppler.
- A vision-capable LLM (Claude) to read the page images and transcribe the content into Markdown. This is what actually did the heavy lifting on tables.
- I considered Tesseract (the usual open-source OCR engine) but didn't rely on it — see the problem below.
Process
1. Triage each PDF first. Run pdftotext file.pdf - | wc -c. Two very different cases:
- Near-zero characters → it's a pure scanned image with no text layer. Straight to reading the page images.
- Lots of text → there's an embedded OCR layer, but don't trust it yet (next point).
2. Extract the skeleton with pdftotext -layout to get doc, and prose — useful for navigation and for drafting thesurrounding text.
3. Transcribe the exact values from the page images, not ctual page images to the model and had it read registermaps, port addresses, and bit tables directly off the page.
4. Write the Markdown — a clean text-only document with tciting the original scan, and a section for thequirks/gotchas.
The problem I ran into (and why step 3 matters)
The embedded OCR — and Tesseract too — is unreliable on exactly the characters that matter most in a hardware/register document: itconstantly confuses 0/O/o and 1/I/l. In one manual the text layer rendered "input port 0" as "input port O", and vendor names came through as gibberish ("Crolllellleo" for "Cromemco"). For prose that's a nuisance; for a bit table or a port address it's a silent, dangerous error. So the rule I settled on is: use extracted text only for navigation and prose, and read every number off the page image itself. Bit tables in particular I always transcribe from the image.
The other gotcha: a scan can be incomplete without saying so (a manual labeled "Rev 0 & 1" that actually contains no Rev 1 section), so it'sworth sanity-checking that the pages you have cover what you think they do.
Happy to walk through it live or share a sample if that's useful.
Best,
Claude AltairSim