The interaction challenges when voice is the primary interface

For decades, human-computer interaction (HCI) was designed around the limits of deterministic software. To use a computer well, you had to make your intent explicit: choose the field, click the button, type the command. Artificial intelligence changes that constraint because computers can now work with partial intent: the messy, unfinished, context-dependent way people actually think and speak.
Voice lets us think aloud, revise ourselves, and omit context. AI lets the system infer what we mean, clarify what’s missing, and remember what we’ve already said. Together, they raise both our expectations and the standard for what we build, while introducing new interaction challenges: making the user experience intuitive, trustworthy, and frictionless.
Speech is an interesting (and challenging) modality
People speak differently from how they type. Speaking often feels more natural, but making sense of it requires the system to fill in missing context and bring structure to free-flowing thoughts. Audio is also noisier than text: microphone quality and background noise affect what the system receives. People are also used to typing, so making voice feel easier means handling that ambiguity for them.
Here are a few decisions we’ve faced:
For the dictation product experience, we had to decide whether to show a live transcript or simply indicate that Flow is listening, then show the final text once someone finishes speaking.
We had to consider a few things:
- Trust: Seeing accurate words appear immediately can help users build confidence in Dictation.
- Accuracy: Waiting for the full utterance gives the model more context. What you say later can clarify what you said earlier.
- Attention: Watching and checking a live transcript pulls attention away from simply speaking.
We also considered other ways to build trust, how quickly we could improve our dictation models, and whether we could show corrections to earlier words as someone continued speaking. Ultimately, we chose non-streaming dictation paired with thoughtful onboarding.
Notetaker summaries involve similar tradeoffs. To reduce cognitive load, the system needs to reliably capture what matters most to each person. With today’s technology, that means balancing summary quality with how long someone has to wait.
Imagine someone in a meeting says, “We should probably move it to next week.” Those words can mean different things depending on who says them and what “it” refers to. Context from Slack or Gmail may also help determine whether the decision matters enough to include in a summary. Gathering that context takes time, as does using more capable models or giving them more time to reason. We have to weigh the improvements in summary quality against the wait.
Handling ambiguity in speech requires choices that span design, ML, product, and engineering. What works for one experience may not work for another, so we revisit those choices for each surface or feature.
New experiences have to earn new habits
Trust in one voice interaction does not automatically transfer to another. Each experience asks our users to adopt a particular behavior, so we have to consider carefully what it takes to build that habit. Take our path from Scratchpad to Notetaker, for example.
We saw users struggling to keep track of quick thoughts while juggling everything else in their day. Some already used voice to capture those thoughts, even returning to their dictation history to recover them. We launched Scratchpad to explore whether we could make that easier.
Dictation competes with typing. When users finish dictating, they expect text to appear immediately and reliably because that is what typing gives them. If they have to wait or keep correcting the output, returning to the keyboard becomes the easier choice. While we’ll continue to push more intelligence into our models, latency and reliability are non-negotiable.
We found that, for many users, building the habit of dictating quick thoughts was challenging, even when they were already comfortable using Flow elsewhere. They often typed into Scratchpad instead. We could have invested more in activation and education, as we had with Flow. But the value of a new experience has to justify the effort of adopting it, and Scratchpad had not yet delivered enough value to earn that new habit.
This forced us to revisit our assumptions. Where could we help users capture and follow up on their thoughts without first asking them to change their behavior?
Meetings presented a similar problem: brief asides and passing suggestions can be easy to miss or hard to keep track of. Crucially, people already talk naturally in meetings. Notetaker lets us build on that behavior and demonstrate value through summaries, MCP, and more. Delivering on that promise requires us to push the frontier of memory and models, building capabilities we’ll need well beyond meetings.
Our goal is to make all of your device interactions delightful, and each step toward that vision needs to deliver enough value to earn the habits it asks you to adopt.
If voice becomes a primary way people interact with computers, there’s an enormous amount left to build. AI opens up new possibilities, but turning them into delightful experiences requires making decisions where few established primitives exist: we have to handle fuzzy intent, choose different tradeoffs for different tasks, and build on habits people can learn to trust.
These are foundational product engineering problems - and a reason there is so much interesting work left to do at Wispr.



.jpg)


