Two weeks ago I wrote about adding voice commands to my daughter’s media player, so a parent in the car can say “[tab name] [song title]” instead of scrolling a phone screen while driving. In that post I said a few tabs were left out on purpose. One of them was a math picture-book tab: 42 books, each with a story audio track and a song. At the time, the reason sounded sensible. The reason didn’t hold up for long.

The Need: The Tab She Asks For Was the One Voice Search Couldn’t Reach

Kids don’t ask for songs by category. My daughter asks for the book she’s thinking about. For the last couple of weeks, a lot of those requests have been math-book songs, and voice search simply didn’t know those books existed. It would say it couldn’t find the song, and we’d be back to the phone screen.

There was a second, smaller problem. When voice search did find something, it only showed a small dark box in the middle of the screen with the title. That’s fine for the driver, who only needs to hear the result. But the person it’s for is sitting in a car seat in the back. For her, a picture says “yes, that’s the one” much faster than a spoken title does.

And the math tab itself had no search box at all. It was a plain list of 42 cards. Finding book 31 meant scrolling.

The Proof: One Prompt at 07:47, Deployed Before 08:00

I log every prompt I type. This is the whole request, translated from Korean:

💬 Prompt that worked “Please make voice search also search the math tab. And when using voice search, can the thumbnail be shown full screen? For the math tab, please add search inside the tab.”

That was at 07:47. The changes were deployed to my NAS and committed before 08:00. What changed:

  • Saying “math [book title]” plays that book’s story audio. Saying “math [book title] song” plays its song instead.
  • When voice search finds a match, the book cover or song picture now fills the whole screen on a black background, with the title in large text at the bottom. Tapping anywhere closes it, and the music keeps playing.
  • The math tab now has a search box that works by book number or title, and it ignores spaces and punctuation.

To be honest about where this stands: I checked the matching logic against all 42 real book titles on my desktop, and every test sentence found the right book and picked story or song correctly. I have not yet used it in the car. Road noise, and whether the full-screen picture actually helps a three-year-old, are still unknown to me.

The Story: “Numbered, Not Named” Was a Guess, Not a Fact

When I excluded that tab two weeks ago, the note in my project file said the titles “weren’t song names” and might collide with other tabs. Looking at it again, that was a guess made in a hurry. The books do have numbers, but every one of them also has a real title, and those titles are what she actually says. A few of them end in odd full-width punctuation like “?” or “!”, which could have caused mismatches, but that turned out to be a one-line fix.

The collision worry was also smaller than I thought. Voice search already narrows to one tab when you start the sentence with a tab name, so “math [title]” never competes with the other tabs.

I think the lesson for me is simple. When I leave something out “for now,” I should write down the exact reason, so I can check later whether the reason is still true. Here it wasn’t, and it took one prompt to fix.

The How: Four Small Changes in One HTML File

The whole app is one index.html file (plain JavaScript, no framework) served by Nginx on a Synology DS224+, with a small Node.js API behind it and a Cloudflare Tunnel in front. Voice recognition uses the browser’s own webkitSpeechRecognition on Android Chrome. Nothing new was added to the stack for this change.

1. Add the tab to the voice keyword list

Voice search checks whether the sentence starts with a known tab name. If it does, it searches only that tab. Adding the math tab was one more line. The real keywords are Korean. These are English stand-ins:

const VOICE_TAB_KEYWORDS = [
  // ...existing tabs...
  { scope: 'math', keys: ['math picture books', 'math'] },  // longer key first
];

Put the longer name first. The code stops at the first key that matches, so if “math” came first, “math picture books” would be cut short.

2. One index entry per book, and decide story or song from the sentence

Before matching, voice search trims filler endings like “play it for me.” One of those endings in Korean is “play the song,” which removes the word “song” from the query. So I don’t check the cleaned-up query for “song”. I check the original sentence instead:

// one entry per book
mathDb.books.forEach(b => {
  idx.push({
    scope: 'math', title: b.title, norm: voiceNormalize(b.title),
    image: `/covers-math/${b.id}.jpg`,
    play: (opts) => playMathAudio(b.id, opts && opts.song ? 'song' : 'story', b.title)
  });
});

// after a match is found
const wantSong = match.scope === 'math'
  && /song/.test(transcript)          // the raw sentence, before fillers were removed
  && !match.norm.includes('song');    // unless "song" is part of the title itself
match.play({ song: wantSong });

The math data is only fetched when someone opens that tab. So before building the index, voice search now loads it if it isn’t loaded yet, the same way it already did for the other lazy tabs.

3. Strip full-width punctuation when comparing titles

Titles and spoken text are compared after removing spaces and punctuation. The old pattern handled ? and ! but not the full-width ? ! , that some Korean book titles use:

function voiceNormalize(s) {
  return (s || '').toLowerCase().replace(/[\s.,!?~\-()\[\]'"·?!,]/g, '');
}

The new search box in the math tab reuses this same function. That’s why typing “birthdayparty” still finds “birthday party”.

4. Full-screen cover on a match

Every entry in the voice index now carries an image URL. Each tab already had one somewhere: book covers for the storybook tabs, the chosen picture for the song tabs, and the cover for the math books. The overlay itself is a plain fixed div:

#voice-cover { display: none; position: fixed; inset: 0; z-index: 260;
               background: #000; flex-direction: column; }
#voice-cover.open { display: flex; }
#voice-cover img { width: 100%; flex: 1; min-height: 0; object-fit: contain; }
function showVoiceCover(url, title, onFail) {
  const cover = document.getElementById('voice-cover');
  const img   = document.getElementById('voice-cover-img');
  document.getElementById('voice-cover-name').textContent = '▶ ' + title;
  img.onload  = () => { hideVoiceOverlay(); cover.classList.add('open'); };
  img.onerror = () => { cover.classList.remove('open'); if (onFail) onFail(); };
  img.src = url;
}

Three choices here were deliberate:

  • It only opens after the image has loaded. If the image fails, the old small text box stays, so there’s never a black screen with nothing on it.
  • It doesn’t close by itself. The small text box used to disappear after about two seconds. A picture for a child in the back seat should stay until someone taps it.
  • Tapping closes the picture, not the music. Starting a new voice search also closes it.

Tabs with no picture, like my plain audio folders, still get the small text box.

One small thing I’d change later: the math covers are stored as 400-pixel-wide copies so the list loads fast. On a large tablet screen that picture may look a little soft. If it bothers her or me, a larger copy just for full screen would fix it.

Cost

$0. No new service, no API key, no new package. Speech recognition and speech output are built into the browser, and everything else is a few dozen lines added to one HTML file on a NAS I already own. The AI work was done inside my existing Claude subscription, with no separate charge for this change.