---
title: "Designing for voice first"
description: "Five principles for designing voice-first interfaces, from making context count to helping people trust what happens when they stop looking at the screen."
date: "2026-10-05T00:00:00.000Z"
canonical: "https://www.antonsten.com/articles/designing-for-voice-first/"
---

Through my work with a startup exploring voice-first interfaces, I’ve noticed how easy it is to design an app around touch and typing, even when the intention is to make it voice-first. You start with familiar components and workflows, and by the time you’re thinking about voice, most of the decisions about how someone will use the product have already been made. Speaking becomes another way to fill in the same fields.

I also caught myself thinking that voice-first meant giving voice, touch, and typing equal importance. People should be able to use whichever makes sense for them, so that felt like a reasonable place to start. But it didn’t really answer what we were trying to explore. If voice is always being fitted into a workflow designed for something else, how much are we actually reconsidering?

I don’t think the answer is to remove the keyboard or make people speak when they’d rather tap. But we do need to question some familiar requirements, like choosing a project before adding a task, or opening the right screen before being able to change something.

These are five principles I’ve been working through. Some get more difficult when you put them together, particularly when you move from someone speaking at their desk to someone talking through their AirPods with their phone in their pocket.

## 1. Make context count

One of the benefits of speaking is how naturally you give more context. If you’ve used voice to prompt an agent, you’ve probably noticed this already. You explain why you’re asking, mention something related, or add a qualification you might not have bothered typing.

For a task, I might type “Review proposal.” Speaking, I’m more likely to say something like, “I need to review the proposal before Friday’s meeting, particularly the pricing, because I’m not sure we’ve allowed enough time for the research.”

There’s useful information in that explanation. The meeting gives the task a timing constraint, and the concern about research tells me what to look for when I come back to it. But if the software turns all of that into a task called “Review proposal,” I haven’t gained much from saying it.

We keep saying that agents work better with more context, but people need a reason to keep providing it. If the extra explanation doesn’t improve the result, we’re effectively teaching them to be brief.

In this example, I’d expect the concern about research to stay with the task and the timing to reflect the meeting. I shouldn’t have to dictate my thought, then open the task and organize everything myself. That’s work the software should be taking on, while leaving me a way to correct its interpretation.

## 2. Let people steer as they think

The proposal example sounds fairly complete written down. In practice, I might say Friday, realize that leaves me no time to prepare, and change it to Thursday while I’m still talking.

That’s something the interaction needs to allow for. People won’t always arrive with a finished instruction, and sometimes seeing the software respond is what helps them figure out what they meant.

If I’m looking at the screen while speaking, I want to see the task taking shape. I can notice that it has the wrong date, add something I forgot, or change direction because the result has made me think of something else. Waiting until I’ve finished talking to show anything loses some of that opportunity.

But showing every partial interpretation could become distracting, too. If the task keeps moving around or changing its wording, I’m now trying to follow the interface while also keeping track of my thought.

I don’t think we’ve resolved all of this yet. The useful question for me is whether the feedback helps someone continue thinking, or gives them another thing to manage.

## 3. Make understanding clear

A microphone animation is useful for knowing that the app is listening, but it doesn’t tell me whether the app has understood what I’m asking.

Even a correct transcript leaves that question open. It might capture both Thursday and Friday perfectly and still assign the wrong one to the task.

When I’m looking at the screen, I can check the result as it appears. Seeing the review scheduled for Thursday, with the concern about pricing attached, tells me more than a generic acknowledgment would. I probably don’t need the app to read it all back as well.

If I’m out walking, that changes. A short spoken response might be useful because I have no other way to know what happened. “Added for Thursday, with a note to check the research budget” gives me enough to catch a misunderstanding without listening to my entire request again.

This is where feedback needs to be specific about what has actually happened. Capturing what I said, understanding it, and acting on it are separate things. If the app has saved my words but hasn’t managed to create the task, I need to know that. Otherwise I’m walking away with confidence it hasn’t earned.

## 4. Make correction easy

Once software starts interpreting what we mean, getting some of it wrong is part of the experience we have to design for.

I should be able to say, “No, the meeting is Friday. I want to review the proposal on Thursday,” and have that fix the relevant detail. I don’t want to repeat the whole request or wonder whether I’ve just created a second task.

There’s also a difference between changing my mind while something is being formed and undoing an action that has already happened. If we’re showing an interpretation before we’re confident in it, the interface needs to make that clear enough that people don’t mistake it for a finished result.

How much confirmation we ask for depends on what the software is about to do. I’m comfortable with a task moving to another day if I can easily move it back. Sending the proposal to a client is a different decision.

Requiring approval for every small change would make voice tedious pretty quickly. But skipping over uncertainty because it makes the demo feel faster won’t hold up in daily use. I need to understand when the system is asking for my judgment and trust that I can repair the ordinary mistakes without starting over.

## 5. Let attention come and go

A lot of what I’ve described so far benefits from being able to see the screen. That’s one of the situations to design for: I’m at my computer or holding my phone, speaking and watching the software respond.

The other is when I’m walking my dog or driving. I want to talk through something without looking at a screen, and I shouldn’t need to take out my phone to confirm a detail or finish the request.

These situations belong in the same product. I might start planning something at my desk and continue thinking about it after I leave. The interaction should survive that change in attention.

That raises some difficult questions about when to interrupt. If something essential is unclear, the app may need to ask me a spoken question. If it can safely keep the thought and leave a detail unresolved, that might be better than breaking my flow. Either way, it needs to avoid leaving me with the impression that something is handled when it isn’t.

When I come back to the screen, I want to see what happened and what still needs me. I don’t want to work through a transcript to discover whether the things I asked for became tasks, remained suggestions, or failed somewhere along the way.

These principles give me something to check my design decisions against, especially when I catch myself reaching for the familiar version of an app again. But this is the part I keep coming back to: seeing the software work can help me trust it, while some of the value of voice comes from being able to stop looking. If I finish a walk and feel the need to check every request, there’s still work to do.
