Conversion from PDF

58 views
Skip to first unread message

Rob Beezer

unread,
Sep 4, 2026, 12:15:16 PMSep 4
to prete...@googlegroups.com, pretext-...@googlegroups.com
(Cross-posting to -dev and -a11y.)

I have (finally) taken back up a project to use AI to convert a PDF into
PreTeXt. I am using examples that have LaTeX source, but that material is not
being used to guide the conversion - only to judge the results. Using
Anthropic's Fable 5.1 right now via Claude Code. I'm hoping to make this
available publicly after a few iterations.

My first test-case is a paper by David Farmer from the arXiv. Many thanks to
David for permission to redistribute the results. Conversion to PreTeXt seems
very successful - much smoother than with ChatGPT about a year ago. Rough
estimates suggest a few minutes per page at a cost of maybe $0.15 per page. And
perhaps that will improve. A little more than a half day of setup and
configuration needed to go from nothing to this first conversion.

One reason for doing this is to take a semi-opaque PDF and make accessible
versions - here PDF, HTML, and braille.

Start at an index of outputs (and the original) at:

https://pretextbook.org/beta/real-roots-2026-09-04/index.html

Some small gotchas in evidence if you look hard enough. And provoking some
places PreTeXt (generally) needs to improve. But I think you have to look hard
for real problems. And there are some surprises - DOI numbers have been
harvested and added to the references (the only allowed content change), and
check out the descriptions on the HTML version of Figure 1.1.

Suggestions and questions are of course welcome.

Rob

David W. Farmer

unread,
Sep 4, 2026, 12:59:04 PMSep 4
to prete...@googlegroups.com

Is there an emoji that is simultaneously a smiley face and
a frowny face? That is how I feel about the descriptions of Figure 1.1.
(It took me many iterations to get those 4 graphs just as I wanted them.)

Regards,

David
> --
> You received this message because you are subscribed to the Google Groups
> "PreTeXt announcements" group.
> To unsubscribe from this group and stop receiving emails from it, send an
> email to pretext-announ...@googlegroups.com.
> To view this discussion visit
> https://groups.google.com/d/msgid/pretext-announce/MTAwMDAzSWgyREZiVzc.1788538513%40pnsh.
>
>

Rob Beezer

unread,
Sep 4, 2026, 1:09:47 PMSep 4
to prete...@googlegroups.com
I agree. :-|

Notable in that Claude just did them without me asking for them, and that they
are not inaccurate (at least a cursory review suggests to me they are not
wrong). Just sort of bland. I didn't look at the figures with ranges marked
off on the horizontal axis.

Is that your only complaint? ;-)

I'm working on a paper of my own with a graph-theory graph, and that is proving
more interesting. Details when that is ready.

Rob

Rob Beezer

unread,
Sep 4, 2026, 1:49:41 PMSep 4
to prete...@googlegroups.com, pretex...@googlegroups.com
OK, one more, and to -a11y on the first try, and not to -announce.

A couple of PreTeXt problems are papered over by hacks in this one, issues at:

https://github.com/PreTeXtBook/pretext/issues/3207

https://github.com/PreTeXtBook/pretext/issues/3208

The graph in Figure 1 is notable. The paper describes it as a Paley graph, with
vertices from the finite field GF(9), and edges based on arithmetic in that
field. So the constructed PreFigure diagram not only copied the visual image
from the paper, but confirmed the vertices and edges with computations and
provides annotations with the actual elements of the field (powers of a
generator for teh non-zero elements). And I guess I said in the published
papaer that the vertices were arranged clockwise, when actually they are
arranged counter-clockwise. Oops. Kinda like posting to the wrong list.

Index at:

https://pretextbook.org/beta/scvt-2026-09-04/index.html

Suggestions on what to hit next are welcome. Needs a PDF, of course. Having
the originating LaTeX then helps with testing the result. And permission to
distribute the derived versions (including PreTeXt source) means we can all have
a look. Papers from the arXiv, with permission, are an obvious choice, but that
will only get us so far. we need new (reasonable) challenges.

I'm off to fix bugs now, I think.

Rob

Rob Beezer

unread,
Sep 7, 2026, 6:41:25 PMSep 7
to prete...@googlegroups.com
An off-list comment from Rick Roesler, which he has allowed me to distribute,
then a follow-up with my comments.

----

If I were going to do this from scratch, I would start with the PreTeXt example
and its pdf. And then build a skill by looping: create pretext from pdf, compare
with the original, update the skill, rinse and repeat. When that works, I would
use some of the other textbook examples where we have the PreTeXt source and the
author-approved pdf to further refine it. I assume you've already done something
like this? I kind of understand why you'd want LaTeX, but it's not clear to me
that it's necessary, and my concern would be that LaTeX dependencies would leak
into the skill and prevent it from working correctly when only pdf was
available. And, by the way, this pdf-to-PreTeXt conversion was one of the skills
I'd thought about from the pretext-authoring plugin.

Rob Beezer

unread,
Sep 7, 2026, 6:42:29 PMSep 7
to prete...@googlegroups.com
Dear Rick,

Yes, I agree your presumed approach makes a lot of sense. Though I had a more
narrow starting point - I wanted to capture existing research articles. Not all
that different, but definitely shorter than a textbook.

When I first started thinking about AI and PreTeXt, I saw a video, and the guy's
point was: AI is great, you can ask it for advice on how to use it! The
converse being, you don't normally fire up Excel and ask it to tell you how to
make a spreadsheet that models cash flow.

Anyway, Claude did not suggest I feed it the PreTeXt Guide nor pretextbook.org,
nor any existing source/PDF pairs. Maybe the reference material is already in
the repo I've been building as an eventual skill. (I haven't looked). I may
release that repo soon.

The LaTeX source is only meant for post-transcoding "scoring". I have asked how
it was used, after-the-fact.

Anyway, I've been spending a lot of time learning (through experience) how I can
accelerate PreTeXt development and adoption, and I have been very happy with the
results. That is not to say I am doing it "right", or that I really have a good
grasp of best practices. And it is like night-and-day compared with a year ago.

I'm going to pursue this a bit further and will then put some results out for
others to look at. Thanks for your two recent issues along these lines - very
interesting, and I have questions, but it could be a few days before I can get
them out.

Thanks,
Rob

Rob Beezer

unread,
Sep 26, 2026, 4:37:24 PM (14 days ago) Sep 26
to prete...@googlegroups.com, pretex...@googlegroups.com
Well, I said I wasn't going to keep posting these, but here goes.

https://arxiv.org/abs/2607.05283, CC-BY

* A nice back story in the New York Times,

https://www.nytimes.com/2026/09/06/science/92-year-old-mathematician-apprentice.html

* Lots of (intricate) diagrams (which was the point). Many cut up to become
#sidebyside, all converted to PreFigure. PDFs of tactile versions, plus,
courtesy of Alexei Kolesnikov, I have about 10 embossed physically, which seem
to have worked out very nicely. (Thanks, Alexei!)

* Transcription, production:
https://pretextbook.org/beta/burau-2026-09-26

Rob

Joseph DiMuro

unread,
Sep 26, 2026, 6:07:38 PM (14 days ago) Sep 26
to bee...@privacyport.com, prete...@googlegroups.com, pretex...@googlegroups.com
Hi Rob. After I'd seen a sample of the diagrams in this paper, I decided to take a look at the PreFigure code for a picture or two.

Okay. That stuff was generated by the AI from the PDF only, right? I am absolutely floored.

Let me remind you of something you said earlier in this thread: "I'm hoping to make this available publicly after a few iterations." Any idea of the timetable for that? I'm not in a rush; I'm just thinking that this might be the thing that gets this AI-hater to start using AI.

It's crazy that I just said that. :-/

-Joseph

--
You received this message because you are subscribed to the Google Groups "PreTeXt accessibility" group.
To unsubscribe from this group and stop receiving emails from it, send an email to pretext-a11y...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/pretext-a11y/MTAwMDAwRmJyVVMuTWM.1790455042%40pnsh.

bee...@privacyport.com

unread,
Sep 26, 2026, 7:30:20 PM (13 days ago) Sep 26
to prete...@googlegroups.com, pretex...@googlegroups.com
Dear Joseph,

Thanks for the reply - responses interspersed.

On 9/26/26 3:07 PM, 'Joseph DiMuro' via PreTeXt development wrote:
> Okay. That stuff was generated by the AI from the PDF only, right?

I believe the graphics were extracted from the PDF as their original PNG
versions, then an AI/Python program "traced" them to PreFigure lines, colors,
regions and splines, etc. I haven't studied the program, but it should be part
of what I release.

> I am absolutely floored.

Me too.

> Let me remind you of something you said earlier in this thread: "I'm hoping to
> make this available publicly after a few iterations." Any idea of the timetable
> for that? I'm not in a rush; I'm just thinking that this might be the thing that
> gets this AI-hater to start using AI.

Let's say I'll get the repository whipped into shape for public hosting sometime
the week of October 5 - I've made myself a note. It won't be finished, but it
might be useful.

> It's crazy that I just said that. :-/

It happens.

Rob

Rob Beezer

unread,
Oct 3, 2026, 4:38:19 PM (7 days ago) Oct 3
to prete...@googlegroups.com, pretex...@googlegroups.com
I just did an experiment with a "random", not too complicated, 12-page paper off
arxiv.org:

https://arxiv.org/abs/2610.02127

First, with the transcription process only looking at the PDF.

Second, allowing the process to take the LaTeX in as input.

I don't have rights to re-distribute, but that is not the point.

The final report is attached. It will be of interest to anybody who has thought
deeply about ambiguities in LaTeX syntax and the struggle to convert LaTeX into
anything else.

A surprise to me: in several instances, Claude wrote "tidier" LaTeX from the
visual PDF rendering, rather than copying what was originally authored (perhaps
due to multiple authors).

Rob
vector-balancing-pdf-only-versus-latex-report-2026-10-03.pdf
Reply all
Reply to author
Forward
0 new messages