This is a follow-up to turning sermon transcripts into a book of verses worth memorizing. That post ended with one finished booklet and an honest worry: would the same script hold up across many more? Two weeks later I had ten booklets, and a new problem. They needed proofreading, and I wanted to hold the result in my hands as one real book.

The Need: Proofreading Is Easier With a Pen

Each booklet is built from a plain-text file. In theory, fixing a typo means opening the text file, finding the sentence, and editing it. In practice, I don’t notice mistakes in a text editor. I notice them when I read the finished page, the same way I’d read a printed book. And I don’t only want to fix typos. I want to say “underline this sentence, it’s the heart of the sermon,” or “this whole entry doesn’t belong, take it out,” or “move this paragraph to the end of the previous entry.”

Those are easy to draw and awkward to type. So the question I started with was simple:

💬 Prompt that worked “I’m proofreading 10 PDFs. How should I request changes, or should I edit them directly?”

The answer shaped everything after it. Don’t edit the PDF itself, because the PDF is regenerated from the text file and any edit there gets overwritten. Mark up the PDF instead. The AI reads the marks and changes the text file. Then the PDF is rebuilt.

The Proof: Ten Booklets Corrected, One Book Printed

Over about a week I marked up all ten booklets in a basic PDF viewer, using its pen and its “add text” tool, and saved each one with _교정 (“corrected”) at the end of the file name. After each one, the AI told me exactly what it had applied, and I checked the rebuilt PDF.

Here’s roughly what went through that loop:

  • Underlines for emphasis: about 115 sentences across ten booklets. A hand-drawn underline became a real, printed underline.
  • Deletions: repeated phrases from the speech-to-text transcript, leftover filler, a few long read-aloud passages.
  • Whole entries removed: five, marked with a big X.
  • Paragraphs moved: one, drawn as a circle with an arrow pointing to where it should go.
  • Verses added or swapped: four, written as a short note like “add Ephesians 4:15” next to the quote box.

Then all ten were bound into one A5 book: 138 pages, with a cover, a table of contents with page numbers, and each sermon starting on a right-hand page.

Total extra cost: $0. Everything ran on my own laptop with free Python libraries (reportlab to build PDFs, PyMuPDF to read and render them) and fonts already installed on Windows. The AI work was part of the Claude subscription I already pay for. I didn’t buy anything new for this.

The Story: The Marks Weren’t Where I Expected

The first surprise came from the very first file. I assumed my pen strokes would be saved as PDF annotations, separate objects a script can list and read. They weren’t. The viewer had flattened them into the page itself, so a script asking “what annotations are on this page?” got zero back.

What worked instead was almost embarrassingly simple. Render each page to a picture and let the AI look at it, the same way a person would look at a marked-up proof. To avoid staring at all 16 pages of every booklet, it first compared the marked-up PDF with the clean one page by page. It counted drawing objects and compared the extracted text, so only pages that had changed got rendered and read.

The marks also turned into a small shared language. I never wrote a style guide. It grew one booklet at a time, and each new mark got written down after the first time I used it:

What I drewWhat it means
A line under wordsPrint this with an underline
Two lines through wordsDelete
A circle, plus a note like A -> BReplace A with B
A big X over a whole entryRemove the entry
A circle and an arrow to another entryMove this text there
A caret with a noteInsert this text here
A note like “add [verse] verse” by the quote boxLook up the verse and add it to the box

After the second booklet I asked one more thing. I had removed an entry and carefully fixed every number after it by hand:

💬 Prompt that worked “Part of entry 6 should move to the end of entry 5, and entry 12 should be deleted. I marked that and also fixed the numbering. Does that make sense? And can the numbering fix itself from now on?”

It can. The PDF builder now ignores whatever numbers are in the text file and counts the entries itself. After that, deleting an entry was just an X.

Some mistakes needed a human decision, and I liked that they came back as questions instead of silent guesses. A speech-to-text error had turned “shattered” into a Korean word that sounds similar but means “exposed.” That was inside a sentence I had underlined but not corrected, so the AI left it alone and asked me. It also misread one of my own instructions. I asked for a subtitle with “br” in it. It treated “br” as an HTML line break and split the subtitle into two lines. It wasn’t a line break, and I had to say so plainly: “don’t drop the br.”

The How

Step 1: Find only the pages I actually marked

import fitz  # PyMuPDF

marked = fitz.open("booklet_marked.pdf")
clean = fitz.open("booklet.pdf")
for i, (a, b) in enumerate(zip(marked, clean)):
    if a.get_text() != b.get_text() or len(a.get_drawings()) != len(b.get_drawings()):
        a.get_pixmap(dpi=110).save(f"marked_p{i+1}.png")   # only these get read

My pen strokes show up as extra drawing objects. Typed notes show up as extra text. Either one flags the page. Small marks, like a strike-through on two words, get a second, zoomed-in render around that spot so the exact range is clear.

🗂 Claude.md Rule A PDF marked up in a basic viewer may save pen strokes flattened into the page, not as annotations — page.annots() returns nothing. Compare against the clean PDF (text + drawing count), render only the changed pages to PNG, and read them visually.

Step 2: Change the source text, never the PDF

Every correction is an exact string replacement in the .txt file, and each one is checked for “this text appears exactly once” before anything is written. If a phrase from my marks doesn’t match the source, nothing is saved and the mismatch is reported. That caught several cases where the words read off the page image didn’t quite match the source, like a comma read as a period, or one word swapped for a similar-looking one. Each time, the AI looked up the exact source text and used that.

An underline is stored in the text file as __like this__ and becomes <u> in the PDF. The default underline sat so close to the letters that it cut through the bottom of some Korean characters, so it was moved down a little: underlineOffset='-0.32*F' in the paragraph style.

Step 3: One cleanup across all ten

Speech-to-text writes numbers as words, so a chapter and verse reference came out spelled in Korean words instead of “12:4”. I asked for all of those to become digits across all ten booklets. A blind find-and-replace is risky here, because many ordinary Korean words start with the same syllables as number words. The AI listed every candidate with its surrounding text, checked each one by meaning, and replaced only the real numbers, together with their context. That came to 45 rules and about 50 places. A second scan afterwards found only non-numbers left.

Step 4: Bind ten into one

A small config file lists the book title, subtitle, a line for the bottom of the cover, the cover image, and the ten text files in order. I put them in date order, from 2018 to 2026. A second script builds the book, reusing all the page styles from the single-booklet script:

  • Cover: a full-page picture for Psalm 119:105, “Thy word is a lamp unto my feet.” The text color is picked automatically from the picture’s brightness: dark text on a light picture, white text with a shadow on a dark one. (How I picked the picture is below.)
  • Table of contents: each sermon title with its year, like “(2018)”, and a page number.
  • Each sermon opens on a right-hand page. The title and the one-page list of verses share that opening page. A separate title page would have been almost empty once I removed the date and speaker line, and would have added ten more pages.
  • Page numbers are centered under the text block. The binding margin is wider on the inside edge and flips between left and right pages. A --oneside option makes every page’s left margin wide, for single-sided printing.

Choosing a cover that doesn’t waste ink

My first plan was to reuse the thumbnail from my kids’ Bible-verse song app, the one for Psalm 119:105: a child carrying a lantern down a path at night, with Jesus walking behind. It looked lovely on screen. On paper, it was a problem. Almost the whole picture is a dark blue night sky, so a color printer would spend most of its ink on the background, on every copy of the cover.

I already had a way around this in another small project of mine, the one where I make coloring pages for my daughter. Its rule is simple: colored outlines, white inside. That way she can color the pages in with crayons, and the white areas use almost no toner. So I ran the same Psalm 119:105 picture through that coloring-page process and looked at the candidate images it gave back. I picked the one that kept only Jesus’ face and the child’s face in full color and left the rest as line art on white. The faces carry the warmth of the original, and the rest of the page stays light.

That choice is also why the cover text color is picked automatically. The first cover was dark and needed white text. The new one is mostly white and needs dark text. Now the script checks the picture and decides.

Step 5: Fix what only shows up in print

These were the three problems I noticed paging through the draft book:

  1. A page with just one line on it. If a paragraph spills one or two lines past the bottom of a page, the line spacing of that paragraph only shrinks, by at most 8%, so it fits. Bigger spills fall back to normal splitting, and allowWidows=0 stops a single last line from going to the next page alone.
  2. The “Pastor’s words” label alone at the bottom of a page, with its paragraph starting on the next page. I fixed this by making the label the first line of the paragraph itself, not a separate heading. reportlab’s orphan control already refuses to leave a paragraph’s first line alone at the bottom, so the label now always travels with its text. Seven cases became zero.
  3. A verse list that spilled one line onto an otherwise empty page. Guessing the height ahead of time didn’t work for longer lists. What worked was building the book, measuring where each list actually ended, and rebuilding with less space above the title if it spilled by 24 mm or less.

🗂 Claude.md Rule To find where something ended up on the page in reportlab, don’t use a zero-height marker flowable. On a full page it gets pushed onto a new page, and the page break after it leaves a blank page. Put the tag on the real paragraph, and pass it on in split() so the second half keeps it too.

The final book: 138 pages. The only blank pages are six intentional ones (behind the cover, behind the table of contents, and before four sermons so they start on the right). No one-line pages, no stranded labels.

The Honest Limit

The cover layout is tuned to one picture. The subtitle was moved down by hand to avoid a small star in the artwork. A different cover image may need the same small adjustment. The line-spacing squeeze is also a trade-off. It’s small enough that I can’t see it on facing pages, but a careful typesetter probably would. And this is still only ten sermons out of a few hundred. The workflow feels settled. The real test is whether the next ten go through it with fewer surprises.


Key Takeaways

  • Don’t edit a generated PDF. Mark it up like paper and fix the source it was built from.
  • A basic PDF viewer may flatten pen marks into the page. Compare with the clean PDF and read only the changed pages as images.
  • A small, growing set of mark meanings (underline, double line, X, arrow, caret) was enough. Write each one down the first time you use it.
  • Let the builder number entries itself, so deleting one is just an X.
  • Print-only problems are fixable: shrink one paragraph’s spacing slightly, keep labels inside their paragraphs, and measure real layout instead of guessing.
  • A picture that looks great on screen can be expensive on paper. A coloring-page style (colored outlines, white inside, faces kept in color) kept the cover warm without a full page of ink.
  • Total extra cost: $0.