A transcript split into labeled speakers — speaker_0 and speaker_1 — with timestamps

Whisper vs ElevenLabs: When a Recording Has More Than One Voice

My wife listens to a Bible teacher every morning. After I built her an app to make that easier (the first post in this series), she asked for one more thing. It turned out to be the hardest part of the project. Could we also read his teachings, she said, not just hear them? Reading helps her study. She wanted the words on a screen. “Sure,” I thought. “I’ll just turn the audio into text.” How hard could it be? In 2026, turning speech into text is mostly a solved problem. ...

June 17, 2026 · 6 min · VibePapa