If you keep an archive long enough, the same thing ends up in it twice. The copies rarely have the same name. They are hard to spot by eye, and even harder once one copy has a transcript and the other doesn’t.
The Need: Two Copies That Don’t Look Like Copies
I run a small, password-protected site at home for one teacher’s recorded lectures. Each episode has audio, video, and (for most of them) a searchable transcript. It lives on my NAS and is sorted into year tabs.
A seven-part lecture series was in there twice:
- One set was in 2017, with titles like “Part (3)” and dates that were all Sundays. These had transcripts.
- The other set was in 2018, with titles like “Lecture 3: …” and dates on random weekdays. These had no transcripts.
The titles did not match, the dates were months apart, and the lengths were different. So was it really the same series, or a second run of the same course with new material?
The Proof: Every Pair Matched, and the Site Now Shows One of Each
I took a 90-second clip from each 2018 episode, turned it into text on my own PC, and compared it to every 2017 transcript. Each clip matched exactly one 2017 episode much more strongly than all the others:
| Lecture | Best match (text overlap) | Next best | Length (2017 / 2018) |
|---|---|---|---|
| 1 | 68% | 20% | 41 / 41 min |
| 2 | 72% | 14% | 45 / 45 min |
| 3 | 40% | 15% | 18 / 56 min |
| 4 | 73% | 24% | 63 / 51 min |
| 5 | 61% | 12% | 24 / 25 min |
| 6 | 71% | 25% | 61 / 44 min |
| 7 | 64% | 24% | 58 / 57 min |
The gap between the best match and the next best was large for every pair. So the answer was yes: all seven lectures were duplicated.
After the cleanup, the site shows one episode per lecture. All seven now have Sunday dates, and six of the seven have a transcript. The public episode count went from 188 to 182. Nothing was deleted. The extra copies are only hidden, so I can bring any of them back.
The Story: I Found It While Making a Booklet
I did not go looking for duplicates. I have been turning some of these lectures into small printed booklets, and I was proofreading one against the site. At one point the transcript window showed an error. While we were looking into that, Claude listed the episodes that were hidden on the site. One of them had almost the same title as a lecture in the 2018 set.
So I asked Claude Code directly:
💬 Prompt that worked “Are there some duplicated ‘soul anatomy’ episodes? Please check.”
That short question was enough. Claude listed every episode with that series name, looked at the audio lengths, and then suggested comparing the actual speech, because the lengths alone could not prove anything.
The part I liked best came after that. I first thought the 2018 dates might be right. Claude pointed out the days of the week. The 2017 dates were all Sundays, which is when these lectures were given. The 2018 dates were a Friday, a Thursday, a Saturday, a Wednesday, and so on. Those look like the days the files were uploaded, not the days the talks were given. Once I saw that, the choice was easy:
“Then 2018.1.7 is right. So make the transcribed episode public and make 2018.4.23 private. Then is the cleanup done?”
For six lectures, I kept the 2017 copy because it had a transcript and the right date. Lecture 3 was the exception. The 2017 copy was only 18 minutes long, and its overlap score was the lowest (40%). My guess is that it was cut short. So for Lecture 3, I kept the 56-minute 2018 copy and moved it to its Sunday date in 2017. That one still needs a transcript.
The How: Clips, Local Speech-to-Text, and Four-Character Overlap
Everything ran on my own PC and NAS. There was no cloud speech API and no extra cost: $0, apart from the Claude Code subscription I already pay for.
1. Cut a short clip from the middle
The first minutes of a talk are often greetings or announcements, which look similar in every episode. So I started each clip at the 5-minute mark and took 90 seconds, in 16 kHz mono (what speech models expect):
ffmpeg -v error -y -ss 300 -t 90 -i "episode.mp3" -ar 16000 -ac 1 clip.wav
2. Turn the clip into text locally
I used faster-whisper with the small model on the CPU. It is not very accurate for Korean, but for this job it doesn’t need to be. It only needs to get enough words right to match.
from faster_whisper import WhisperModel
model = WhisperModel("small", device="cpu", compute_type="int8")
segments, _ = model.transcribe("clip.wav", language="ko", beam_size=1)
text = " ".join(s.text for s in segments)
Seven clips took a few minutes in total.
3. Compare with four-character pieces, not words
Speech-to-text often gets Korean word spacing wrong, so comparing whole words fails. Instead, I removed all spaces and cut the text into overlapping 4-character pieces. The score is how many of the clip’s pieces also appear somewhere in a transcript:
import re
def grams(s):
s = re.sub(r"\s+", "", s)
return {s[i:i+4] for i in range(len(s) - 3)}
clip = grams(text)
score = len(clip & grams(transcript)) / len(clip)
A real match came out between 40% and 73%. A non-match, from the same speaker using the same kinds of words, came out between 12% and 25%. What matters is the gap between the two groups, not the exact numbers.
4. Hide, don’t delete
Each episode in the site’s data has a published flag. Hiding a duplicate is just "published": false. The audio, video, and transcript files stay on the NAS. For the one episode that moved from 2018 to 2017, I also changed its year, its ID, and the date inside its file names, and then checked that the audio and video still played on the live site.
A mistake along the way
Before we found the duplicates, I asked how many episodes there were. Claude counted every entry in the data files and got 191, while the site’s index said 188. I asked it to “fix” the index to 191, and it did.
It was wrong. When Claude later compared the local copy with the NAS, it found three episodes marked published: false and checked how the server builds the index. The server only counts published episodes, so 188 had been correct from the start. Claude told me this directly, put the index back to 188, and wrote the rule down so it would not happen again:
🗂 Claude.md Rule
index.jsoncounts only episodes withpublished: true(the server’ssyncIndex). If some episodes are hidden, the index total will be smaller than the number of entries in the episode files. That is normal. Do not “fix” it by hand.
The mistake was cheap to undo. It was also useful, because the three hidden episodes it found were what made me look more closely at the series.
TL;DR
- Titles, dates, and lengths could not prove the two sets were the same. Comparing the actual speech could.
- A 90-second clip from the middle, local speech-to-text, and 4-character overlap was enough. The best match was always far above the rest.
- The day of the week showed which dates were real (Sunday lectures) and which were upload dates.
- Hide duplicates with a flag instead of deleting them.
- Before correcting a total, check how the system calculates it.
- Cost: $0 beyond an existing subscription. Everything ran locally.