▶  "Call mom"  →  "Calling Tom..."

▶  4-yr-old: "Play Elsa"  →  "Playing Alexa radio"

▶  British accent: "tomato"  →  [confused silence]

Try itClick each ▶ to see what the assistant did.

Today: why this is one of the
hardest problems in AI.

IEEE STEM SUMMIT 2026

How Voice Assistants
Hear You

The hidden science of speech AI
DEEPAK PISKALA  ·  19 YEARS IN AI  ·  AUTHOR, BUILDING SPEECH AI
02
The question this booth answers
How does a microphone turn the words
"play despacito" into Despacito actually playing?
03
Part 1 · What the machine actually sees

Your voice is just a wiggle.

A microphone records pressure changes ~16,000 times per second.

That's all your voice assistant gets. A wiggle.
04
Part 2 · Sound is made of Lego bricks

English has ~44 sounds. Every word is built from them.

Try itClick a brick in the top row to swap its sound.

/k/
/æ/
/t/
→ CAT

/d/
/ɔ/
/g/
→ DOG

These tiny units are called phonemes. The first job of a voice assistant: listen to the wiggle and guess which Legos.

05
Part 3 · Why this is hard

These two sentences sound identical.

"RECOGNIZE SPEECH"
"WRECK A NICE BEACH"

↑ Same waveform. You know which is which because you know the conversation. Machines are still catching up.

Try itWhat was the conversation about?

06
Part 4 · The anatomy of a voice assistant

Inside Alexa, Siri, Google — the same 7-stage pipeline.

From the moment you speak the wake word to the moment the device speaks back, your voice passes through seven specialised AI systems — each one a research field of its own.
🗣️
YOU
→
1
Wake Word
"Alexa", "Hey Siri" — always listening, on-device, ultra low power.
→
2
ASR
Speech → text. The "wiggle to words" model.
→
3
NLU
What did you mean? Intent + slots.
→
4
Dialog Mgr
Track context. Ask follow-ups. Remember turn-to-turn.
→
5
Skill / Action
The actual thing: play music, set timer, call mom.
→
6
NLG
Write a natural reply ("Playing Despacito").
→
7
TTS
Text → speech. The voice you hear.
→
🔊
REPLY
Try it

⚡ Why this matters

If any one of these 7 stages misfires, the whole experience breaks. ASR mishears "mom" as "Tom" → you call the wrong person. That's why voice AI is fragile even after decades of research.

🎓 Each box = a STEM career

Signal processing (wake word), deep learning (ASR), linguistics (NLU), reinforcement learning (dialog), software engineering (skills), generative AI (NLG), audio synthesis (TTS).

07
Part 5 · How we got here

For 30+ years, voice AI was many small models stitched together.

RULES ERA1950s – 1970s

"If it sounds like this, then it's that."

Engineers hand-built rules and sound templates. Brittle. Worked for a tiny vocabulary in a quiet room (Bell Labs' 1952 "Audrey" knew only spoken digits). Failed the moment you had a cold.

PIPELINE ML1980s – mid-2010s

A factory line of small ML brains, each doing one job.

Acoustic model
→
Pronunciation
→
Language model
→
Decoder
→
Text out

Each box was its own model, trained separately. If one was weak, the whole pipeline got worse. Errors compounded.

08
Part 5 (cont.) · The breakthrough

Today: one giant brain learns the whole job.

audio(the wiggle)
→
END-TO-END
NEURAL MODEL~1 billion parameters
→
text(what you said)
No hand-written rules. No separate pipeline stages. The model figures it out from millions of examples.
That shift, plus far more training data, is why today's recognizers handle accents, noise, and casual speech much better than a decade ago.
09
⚡ Try it yourself — 3 minutes

Let's break it on purpose.

Open the voice-typing feature on a phone or Chromebook (or any free browser-based recognizer) and try to fool it three ways.

🤫

The Whisperer

Speak a sentence at half-volume. Watch which words go missing.

🌍

The Accent

Same sentence — different regional accent. Note which words slip.

🎵

The Noisy Room

Speak with background music. Watch the model hallucinate lyrics.

↳ Try it live with me in this booth: Fri 23 Oct, 12:45–1:30 PM ET↳ Then score it: paste what it wrote into the Be the Recognizer scorer
10
Why this matters in your classroom

Voice assistants are worse at understanding the people in your room.

2-3×
higher error rate
on children's speech
vs. adults.
📏

Shorter vocal tracts

Higher pitch, different formants — most models weren't trained on them.

👅

Developing pronunciation

"R", "S", "Th" come online at different ages.

📊

Less training data

Privacy rules limit children's audio in datasets.

Try it

Your students are the underserved users — and the future researchers who fix this.
Source: The Learning Agency, "Closing the Child Speech Recognition Gap": about 5% word error rate on adults vs. 11–18% on children for the same models.
11
Part 6 · What's next

The future isn't voice alone. It's multimodal.

When you talk to a friend, they don't just hear you. They see your face, follow your gaze, watch your hands, remember the conversation. The next generation of AI does the same.

Try itYou say “Put that there.” Switch on more senses above.
Inspired by MIT’s 1980 “Put-That-There” demo.
▼ fused inside one model ▼
A truly conversational AI understands you the way a person would.
12
Open problems = tomorrow's careers

Six unsolved problems waiting for your students.

🌐
LINGUISTICS + ML

Low-resource languages

7,000+ human languages. Only a small fraction have enough data for voice AI to work well. Whole communities are locked out.

🔀
SOCIO-LINGUISTICS

Code-switching

Bilingual speakers mix languages mid-sentence (Spanglish, Hinglish). Models break. Humans don't.

❤️
AFFECTIVE COMPUTING

Emotion & intent

Sarcasm. Urgency. Distress. Same words, different meaning. Today's AI hears the words and misses the human.

👶
DEVELOPMENTAL AI

Voices for everyone

Children, elderly, speech disorders. Build assistants that work for the people they currently fail.

🔒
PRIVACY + EDGE ML

On-device recognition

Run capable speech models on small, low-power chips so audio never leaves the device. Privacy by default.

🎭
MULTIMODAL FUSION

Voice + vision + gesture

Combine signals the way humans do. Still wide open for new ideas.

📣 Every one of these is a PhD thesis, a startup, and a 30-year career — waiting for a student you teach next semester.
13
Take this back to your classroom — 15 min, no equipment

Activity: "Be the Recognizer"

1

Pair students up

One speaker, one "recognizer" (writes down what they hear).

2

Add noise

Round 1 in a quiet room. Round 2 with music playing while the speaker whispers.

3

Tally errors

Compare what was said to what was written. Calculate the class-wide "word error rate".

4

Discuss

What made it hard? Where would context have helped? Bridge directly to how AI faces the same problem.

📄

FREE WORKSHEET + TEACHER GUIDE
Download the PDF from this booth's Resources (printable, no logins)Play it online, or print it from this kit

14
For you, for your classroom — all free

Resources

🧪

Teachable Machine (Google)

Train a voice classifier in a browser in 10 minutes. No code.

Open ↗
🤗

Hugging Face Spaces

Hundreds of free speech-AI demos to try in a browser. Preview before class.

Open ↗
📘

Building Speech AI

Free companion code at github.com/prdeepakbabu/building-speech-ai

Open ↗
🛠️

IEEE TryEngineering

Free, ready-to-use engineering lesson plans for pre-university classrooms.

Open ↗
15
The next breakthrough in voice AI
is sitting in your classroom.
DEEPAK PISKALA  ·  @prdeepakbabu  ·  prdeepakbabu.github.io