Conversion from PDF

29 views
Skip to first unread message

Rob Beezer

unread,
Sep 4, 2026, 12:15:16 PMSep 4
to prete...@googlegroups.com, pretext-...@googlegroups.com
(Cross-posting to -dev and -a11y.)

I have (finally) taken back up a project to use AI to convert a PDF into
PreTeXt. I am using examples that have LaTeX source, but that material is not
being used to guide the conversion - only to judge the results. Using
Anthropic's Fable 5.1 right now via Claude Code. I'm hoping to make this
available publicly after a few iterations.

My first test-case is a paper by David Farmer from the arXiv. Many thanks to
David for permission to redistribute the results. Conversion to PreTeXt seems
very successful - much smoother than with ChatGPT about a year ago. Rough
estimates suggest a few minutes per page at a cost of maybe $0.15 per page. And
perhaps that will improve. A little more than a half day of setup and
configuration needed to go from nothing to this first conversion.

One reason for doing this is to take a semi-opaque PDF and make accessible
versions - here PDF, HTML, and braille.

Start at an index of outputs (and the original) at:

https://pretextbook.org/beta/real-roots-2026-09-04/index.html

Some small gotchas in evidence if you look hard enough. And provoking some
places PreTeXt (generally) needs to improve. But I think you have to look hard
for real problems. And there are some surprises - DOI numbers have been
harvested and added to the references (the only allowed content change), and
check out the descriptions on the HTML version of Figure 1.1.

Suggestions and questions are of course welcome.

Rob

David W. Farmer

unread,
Sep 4, 2026, 12:59:04 PMSep 4
to prete...@googlegroups.com

Is there an emoji that is simultaneously a smiley face and
a frowny face? That is how I feel about the descriptions of Figure 1.1.
(It took me many iterations to get those 4 graphs just as I wanted them.)

Regards,

David
> --
> You received this message because you are subscribed to the Google Groups
> "PreTeXt announcements" group.
> To unsubscribe from this group and stop receiving emails from it, send an
> email to pretext-announ...@googlegroups.com.
> To view this discussion visit
> https://groups.google.com/d/msgid/pretext-announce/MTAwMDAzSWgyREZiVzc.1788538513%40pnsh.
>
>

Rob Beezer

unread,
Sep 4, 2026, 1:09:47 PMSep 4
to prete...@googlegroups.com
I agree. :-|

Notable in that Claude just did them without me asking for them, and that they
are not inaccurate (at least a cursory review suggests to me they are not
wrong). Just sort of bland. I didn't look at the figures with ranges marked
off on the horizontal axis.

Is that your only complaint? ;-)

I'm working on a paper of my own with a graph-theory graph, and that is proving
more interesting. Details when that is ready.

Rob

Rob Beezer

unread,
Sep 4, 2026, 1:49:41 PMSep 4
to prete...@googlegroups.com, pretex...@googlegroups.com
OK, one more, and to -a11y on the first try, and not to -announce.

A couple of PreTeXt problems are papered over by hacks in this one, issues at:

https://github.com/PreTeXtBook/pretext/issues/3207

https://github.com/PreTeXtBook/pretext/issues/3208

The graph in Figure 1 is notable. The paper describes it as a Paley graph, with
vertices from the finite field GF(9), and edges based on arithmetic in that
field. So the constructed PreFigure diagram not only copied the visual image
from the paper, but confirmed the vertices and edges with computations and
provides annotations with the actual elements of the field (powers of a
generator for teh non-zero elements). And I guess I said in the published
papaer that the vertices were arranged clockwise, when actually they are
arranged counter-clockwise. Oops. Kinda like posting to the wrong list.

Index at:

https://pretextbook.org/beta/scvt-2026-09-04/index.html

Suggestions on what to hit next are welcome. Needs a PDF, of course. Having
the originating LaTeX then helps with testing the result. And permission to
distribute the derived versions (including PreTeXt source) means we can all have
a look. Papers from the arXiv, with permission, are an obvious choice, but that
will only get us so far. we need new (reasonable) challenges.

I'm off to fix bugs now, I think.

Rob

Rob Beezer

unread,
Sep 7, 2026, 6:41:25 PM (11 days ago) Sep 7
to prete...@googlegroups.com
An off-list comment from Rick Roesler, which he has allowed me to distribute,
then a follow-up with my comments.

----

If I were going to do this from scratch, I would start with the PreTeXt example
and its pdf. And then build a skill by looping: create pretext from pdf, compare
with the original, update the skill, rinse and repeat. When that works, I would
use some of the other textbook examples where we have the PreTeXt source and the
author-approved pdf to further refine it. I assume you've already done something
like this? I kind of understand why you'd want LaTeX, but it's not clear to me
that it's necessary, and my concern would be that LaTeX dependencies would leak
into the skill and prevent it from working correctly when only pdf was
available. And, by the way, this pdf-to-PreTeXt conversion was one of the skills
I'd thought about from the pretext-authoring plugin.

Rob Beezer

unread,
Sep 7, 2026, 6:42:29 PM (11 days ago) Sep 7
to prete...@googlegroups.com
Dear Rick,

Yes, I agree your presumed approach makes a lot of sense. Though I had a more
narrow starting point - I wanted to capture existing research articles. Not all
that different, but definitely shorter than a textbook.

When I first started thinking about AI and PreTeXt, I saw a video, and the guy's
point was: AI is great, you can ask it for advice on how to use it! The
converse being, you don't normally fire up Excel and ask it to tell you how to
make a spreadsheet that models cash flow.

Anyway, Claude did not suggest I feed it the PreTeXt Guide nor pretextbook.org,
nor any existing source/PDF pairs. Maybe the reference material is already in
the repo I've been building as an eventual skill. (I haven't looked). I may
release that repo soon.

The LaTeX source is only meant for post-transcoding "scoring". I have asked how
it was used, after-the-fact.

Anyway, I've been spending a lot of time learning (through experience) how I can
accelerate PreTeXt development and adoption, and I have been very happy with the
results. That is not to say I am doing it "right", or that I really have a good
grasp of best practices. And it is like night-and-day compared with a year ago.

I'm going to pursue this a bit further and will then put some results out for
others to look at. Thanks for your two recent issues along these lines - very
interesting, and I have questions, but it could be a few days before I can get
them out.

Thanks,
Rob
Reply all
Reply to author
Forward
0 new messages