Joseph DiMuro would like to experiment with the PDF
transcription process, so I am going to make that available
now. This all assumes you have Claude Code, and can run
"claude" at the terminal. (This is not a chatbot-style
interaction.)
* Copy or clone
https://github.com/PreTeXtBook/pdf-to-pretext
* In either case, *copy* the skill/pdf-to-pretext
directory to ~/.claude/skills/pdf-to-pretext . Claude
will then find the skill from any directory where you have
a PDF to transcribe.
* Or symlink ~/.claude/skills/pdf-to-pretext to that
directory in a clone.
* Note: you can update a clone easily with a "git pull".
Copies will need manual updates. See below about pull
requests, which will also require a clone.
* Then you just tell Claude something like: "Use the
pdf-to-pretext skill to transcribe foo.pdf into PreTeXt.
foo.tex is/is not available as additional input."
On first use, Claude will get a clone of PreTeXt and build a
Python virtual environment with a pile of tools, for running
pretext/pretext - the skill knows what needs to be done.
You do need a TeX installation with xelatex, and it will
give you the command to install one if it is missing.
The result is a PreTeXt project, built to HTML and PDF. The
new PDF is compared to the original: on four papers so far,
96 to 98 percent of the words match (not 100, since the
bibliography gets reformatted and mathematics does not
extract as words). Expect a quarter-hour to an hour per
paper.
The skill keeps notes on what it had to figure out for
itself, and will offer to draft an issue from them. Please
send those along. It is supposed to leave out the quirks of
your one paper, and a pull request should too - a rule
learned from one paper is usually wrong for the next ten.
See CONTRIBUTING.md.
Rob