Be the Recognizer

Today you are the voice assistant. Can you beat the noise? Listen, write what you hear, then score yourself the way engineers score speech recognizers. About 15 minutes, in pairs, grades 4–12.

🔒 Nothing is recorded, saved or sent. What you type stays on this page.

How to play

1 · Partner up

The Speaker controls the cards below. The Recognizer must not peek.

2 · Round 1: quiet

The Speaker reads cards 1–3 once each, in a normal voice.

3 · Round 2: noisy

Turn on noise or music. The Speaker whispers cards 4–6, once each.

4 · Score it

The Recognizer types what they heard below, and the page counts the errors.

One device per pair? The Speaker reads from the screen while the Recognizer writes on paper, then types it in to score. Two devices work too: one for the cards, one for scoring.

1. Speaker cards (Speaker only!)

Card hiddenTurn the screen away from the Recognizer, then press Show card.

Speaker rules: read each card one time. No repeats, no spelling, no hints. A real voice assistant gets one try, too.

2. Score it

Word error rate = (swapped + missing + extra) ÷ words said saidheard swapped said missing +heard extra

Type exactly what you heard, even if it makes no sense, then press Score. Heard nothing? Leave it blank and press Score. Capitals and punctuation don’t count, and “4” matches “four”.

3. Score any sentence

For your own cards, or for human vs. machine: read a card to a voice-typing tool and paste what it wrote. The worked example below scores 75%.

4. Think about it

1. Which round had the higher word error rate? What did the noise and whispering take away?

2. When did the other words in the sentence help you guess right, or trick you?

For teachers

At a glance

15 minutes · grades 4–12 · pairs, or trios with a scorekeeper · one device per pair, or the printable version (recognizer sheet, speaker cards to cut out, teacher guide).

Run of show

  1. Hook (2 min). “Has a voice assistant ever misheard you?” Today, students are the voice assistant.
  2. Round 1 (3 min). Quiet room. Speakers read cards 1–3 once, in a normal voice.
  3. Round 2 (3 min). Play music or the noise button at a moderate volume. Speakers whisper cards 4–6 once.
  4. Score (4 min). Recognizers type what they heard. Collect each pair’s round results; class WER = all errors ÷ all words, for each round.
  5. Discuss (3 min). Use the prompts below. Bonus cards are for early finishers.

Scoring notes

  • The scorer ignores capitals and punctuation and treats numerals like “4” and “13th” as words. Engineers call this text normalization.
  • Sound-alikes with a different spelling (for/four, pears/pairs, two/to) are errors: a machine that writes the wrong word may do the wrong thing.
  • WER can exceed 100% when many extra words are written.

Discussion prompts

  • “I scream” or “ice cream”? The sounds are nearly the same. How did you decide?
  • Why was “Siobhan” so hard? You have probably heard that name far less often than “Sarah”.
  • Who might voice assistants mishear most often? How could engineers fix that?

The science connection

  • Speech recognizers are graded with this exact metric: word error rate, WER = (S + D + I) ÷ N.
  • A recognizer combines an acoustic model (which sounds did I hear?) with a language model (which words usually go together?). Guessing “ice cream” from context is the language model at work.
  • Noise and whispering hide acoustic detail. Engineers train models on audio mixed with noise so they learn to cope.
  • Rare words and names remain hard for people and machines alike, and improving them is an active research area.

What each card tests

1Sound-alikes: Mom / Tom, four / for
2Word boundaries: “I scream” / “ice cream”
3Numbers that sound alike: fifteen / fifty
4Word boundaries: “recognize speech” / “wreck a nice beach”
5Homophones: eight / ate, pears / pairs
6A rare name (Siobhan) plus two / to / too
7Word boundaries: “stuffy nose” / “stuff he knows”
8Numbers that sound alike: thirteenth / thirtieth

Adaptations and extensions

  • Younger students: skip the formula and count “oops words” per round. Grades 9–12: explain why WER can exceed 100%, and graph class results by round.
  • Human vs. machine: read the same cards to a voice-typing tool and paste its text into Score any sentence. Many online speech tools send audio to the cloud, so follow your school’s technology and privacy policy.
  • See the noise: open See Your Voice and whisper a card: the voice’s buzz stripes disappear.
  • AI4K12 connection: Big Idea 1 (Perception) and Big Idea 4 (Natural Interaction).

Privacy

No sign-in or ads. Basic page analytics records visits, while activity answers stay in this tab. The activity never uses the microphone; the noise button only plays sound, and typed answers disappear when the tab closes.