Today: why this is one of the hardest problems in AI.
IEEE STEM SUMMIT 2026
How Voice Assistants Hear You
The hidden science of speech AI
DEEPAK PISKALA · 19 YEARS IN AI · AUTHOR, BUILDING SPEECH AI
02
The question this booth answers
How does a microphone turn the words "play despacito" into Despacito actually playing?
03
Part 1 · What the machine actually sees
Your voice is just a wiggle.
A microphone records pressure changes ~16,000 times per second.
That's all your voice assistant gets. A wiggle.
04
Part 2 · Sound is made of Lego bricks
English has ~44 sounds. Every word is built from them.
Try itClick a brick in the top row to swap its sound.
/k/
/æ/
/t/
→CAT
/d/
/ɔ/
/g/
→DOG
These tiny units are called phonemes. The first job of a voice assistant: listen to the wiggle and guess which Legos.
05
Part 3 · Why this is hard
These two sentences sound identical.
"RECOGNIZE SPEECH"
"WRECK A NICE BEACH"
↑ Same waveform. You know which is which because you know the conversation. Machines are still catching up.
Try itWhat was the conversation about?
06
Part 4 · The anatomy of a voice assistant
Inside Alexa, Siri, Google — the same 7-stage pipeline.
From the moment you speak the wake word to the moment the device speaks back, your voice passes through seven specialised AI systems — each one a research field of its own.
The actual thing: play music, set timer, call mom.
→
6
NLG
Write a natural reply ("Playing Despacito").
→
7
TTS
Text → speech. The voice you hear.
→
🔊
REPLY
Try it
⚡ Why this matters
If any one of these 7 stages misfires, the whole experience breaks. ASR mishears "mom" as "Tom" → you call the wrong person. That's why voice AI is fragile even after decades of research.
🎓 Each box = a STEM career
Signal processing (wake word), deep learning (ASR), linguistics (NLU), reinforcement learning (dialog), software engineering (skills), generative AI (NLG), audio synthesis (TTS).
07
Part 5 · How we got here
For 30+ years, voice AI was many small models stitched together.
RULES ERA1950s – 1970s
"If it sounds like this, then it's that."
Engineers hand-built rules and sound templates. Brittle. Worked for a tiny vocabulary in a quiet room (Bell Labs' 1952 "Audrey" knew only spoken digits). Failed the moment you had a cold.
PIPELINE ML1980s – mid-2010s
A factory line of small ML brains, each doing one job.
Acoustic model
→
Pronunciation
→
Language model
→
Decoder
→
Text out
Each box was its own model, trained separately. If one was weak, the whole pipeline got worse. Errors compounded.
08
Part 5 (cont.) · The breakthrough
Today: one giant brain learns the whole job.
audio(the wiggle)
→
END-TO-END NEURAL MODEL~1 billion parameters
→
text(what you said)
No hand-written rules. No separate pipeline stages. The model figures it out from millions of examples.
That shift, plus far more training data, is why today's recognizers handle accents, noise, and casual speech much better than a decade ago.
09
⚡ Try it yourself — 3 minutes
Let's break it on purpose.
Open the voice-typing feature on a phone or Chromebook (or any free browser-based recognizer) and try to fool it three ways.
🤫
The Whisperer
Speak a sentence at half-volume. Watch which words go missing.
🌍
The Accent
Same sentence — different regional accent. Note which words slip.
🎵
The Noisy Room
Speak with background music. Watch the model hallucinate lyrics.
Voice assistants are worse at understanding the people in your room.
2-3×
higher error rate on children's speech vs. adults.
📏
Shorter vocal tracts
Higher pitch, different formants — most models weren't trained on them.
👅
Developing pronunciation
"R", "S", "Th" come online at different ages.
📊
Less training data
Privacy rules limit children's audio in datasets.
Try it
Your students are the underserved users — and the future researchers who fix this.
Source: The Learning Agency, "Closing the Child Speech Recognition Gap": about 5% word error rate on adults vs. 11–18% on children for the same models.
11
Part 6 · What's next
The future isn't voice alone. It's multimodal.
When you talk to a friend, they don't just hear you. They see your face, follow your gaze, watch your hands, remember the conversation. The next generation of AI does the same.
🎤
Voice
What you said
👁️
Vision
Lip-reading, gaze, who's in the room
✋
Gesture
Pointing, nodding, sign language
🧠
Context
What you said 5 minutes ago, where you are
Try itYou say “Put that there.” Switch on more senses above.
Inspired by MIT’s 1980 “Put-That-There” demo.
▼ fused inside one model ▼
A truly conversational AI understands you the way a person would.
12
Open problems = tomorrow's careers
Six unsolved problems waiting for your students.
🌐
LINGUISTICS + ML
Low-resource languages
7,000+ human languages. Only a small fraction have enough data for voice AI to work well. Whole communities are locked out.