Skip to Content
0%

Six Principles of Voice Design: Designing for the Ear and Encoding it into Agent Script

Play before-and-after audio examples of six voice design principles, then see how to encode each one in Agent Script.

“Hello? Are you still there?” asks the caller in the first audio example below, after four seconds of silence from a voice agent. Each of the six voice design principles here comes with audible before-and-after examples like that one, plus the Agent Script that encodes the fix.

Press play as you read.


Table of Contents


Introduction: Why voice design needs its own rules

The voice channel has long occupied an awkward position in the enterprise stack: perpetually promising, rarely delivering. Voice AI struggled mostly because the technology fell short of consumer expectations, with poor natural language understanding, no contextual memory, and fragile multi-turn handling. Large (and small) language models have gone a long way toward overcoming these limitations. Enterprise voice AI adoption is now an operational reality for a growing share of the market,[1] and it’s coming for most of the rest. Technology will keep improving, and those improvements will keep mattering for user experience and business value, but technology is no longer the bottleneck.

The biggest challenge now is designing the user interface. Designing for the ear is fundamentally different from designing for the eye. You can’t lift a chatbot onto a phone line and call it a voice agent.

Voice users behave differently. They are more sensitive to latency. They notice instantly when an agent mispronounces something or hesitates in the wrong place. The conversational flow is more fluid than chat, with more turns, repair, overlap, disfluencies, cut-offs, and subtle variation. None of that is informational noise.

Voice also removes everything chat gives you visually: no buttons, no links, no forms, no images, no tables, no visual hierarchy.

You can teach users where to click. You can’t teach them a new way to talk.

For engineers, architects, admins, and designers, this guide outlines core principles for voice user interface (VUI) design, along with the background you need to build agents that deliver naturalistic real-time voice experiences. Its human-centered design approach draws on conversation analysis (CA), the study of how turn-by-turn talk is organized. It also draws on what we know about human interactional-linguistic behavior and cognitive limits.

In human-computer interaction, “conversational” is often taken to mean persona and tone of voice, and conversation design is sometimes reduced to just that. In conversation analysis, “conversational” points instead at the systematics and mechanics of social interaction (Stokoe et al., 2024). Sixty years of research in CA show the remarkable systematicity of talk-in-interaction, despite the folk platitude that conversation is “messy.”

Design that runs against these systematics gets punished for it: the interface feels off in ways callers can’t articulate, and adoption drops. Fortunately, those conversational rules are extensively documented. The sections below apply that research to voice experiences and show how to encode each recommendation in Agentforce’s Agent Script.[2]

🔊 Each principle includes before-and-after audio examples you can play.

Three layers of voice design in Agentforce

Designers and engineers can’t invent the rules of conversation, but they can encode designs that respect them. In Agentforce, voice design spans three distinct layers:

LayerPurpose
1. Prose prompt instructions (|)Controls conversational behavior: e.g., turn shape, end focus, repair phrasing. Written in Agent Script as concrete turn rules (e.g., “Place the question at the very end”), not vague tone descriptors (e.g., “Be helpful”).
2. Deterministic logic & variablesGuarantees sequence: e.g., slot-filling variables, retry counters, mandatory summary confirmations, automated escalation logic.

Two blocks, before_reasoning 

and after_reasoning, run outside the reasoning loop for guarantees that must hold every turn.
3. Voice channel configurationHandles sound and timing: e.g., text-to-speech (TTS) voice selection, custom pronunciations, filler turns (messaging and audio), acoustic barge-in, telephony routing.
Authored in a top-level modality voice: block; the compiler enforces ranges on timing values.

Each of the six principles below ends with its encoding, split across these same three layers.

Choosing a voice

The voice itself is one line of script, but it’s a persona decision before it’s a configuration decision. Pick it against the agent’s persona spec, not from whatever the library offers first (Lucy, 2026).[3] Mechanically, voice_id is an org-scoped identifier from your org’s voice library, picked in Agentforce Builder under Voice Settings, which also holds pronunciation dictionaries and key-term prompting. It doesn’t carry over to other orgs.

There is also a speech model behind the voice. Fast, natural-sounding speech is becoming a commodity. Providers such as ElevenLabs illustrate how far text-to-speech has come. Still, a human-sounding voice is only one part of a good voice agent. An agent can sound remarkably human and still leave a caller in silence, ask three questions at once, or keep talking after the caller has taken the floor.

It helps to separate the voice from the model that produces it, and both from the agent’s conversational behavior. In Agentforce, voice selection and speech-model configuration belong in the voice channel layer. The available models, voices, and delivery controls will change. Design around the language, pronunciation, timing, and persona your callers need, not around a particular provider’s settings.

One timing value in the block below matters for Principle 1: beepboop_config sets how long a backend call runs before the platform’s filler audio starts.

NOTE
Code blocks throughout are excerpts. A deployable script also needs a config: block and exactly one start_agent, and every subagent needs a description:. Those are omitted here for brevity.

Agent Script: Voice channel configuration

connection telephony:
  modality voice:
    voice_id: "vX7kQm2LpRa9TnBcYdEf"   # org-scoped ID from Builder > Voice Settings
    outbound_speed: 0.95
    outbound_stability: 0.85
    outbound_similarity: 0.85
    additional_configs:
      beepboop_config:
        max_wait_time_ms: 500          # audio cue at 0.5s to mask processing delay

When the default speech model doesn’t meet your callers’ needs, Agent Script also lets you select a supported model and a compatible voice. The example below uses an ElevenLabs model to show that separation. Delivery controls depend on the selected model. See Agentforce’s voice-model reference for the current options and configuration details.

Agent Script: Selecting a voice model

# One supported model and a compatible voice, illustrative
modality voice:
  outbound:
    persona_id: "a029af9d692c"     # voice ID (model-specific)
    model:
      id: "eleven_flash_v3"

Most of a voice agent’s value comes from orchestration and interaction design. That means connecting the caller’s request to the right action, managing the wait, and shaping each turn so the caller knows what happens next. A natural-sounding voice matters, but it can’t compensate for an interaction that runs against how conversation works. That’s the work of the six principles below.


Designing for the ear: Six principles of voice design

Successful voice user interface design must respect the systematics of conversation and the limits of human cognitive processing. The six principles below are fundamental to VUI design. Each comes with concrete examples. They’re a starting point rather than an exhaustive guide.[4]

Example scenarios are based on forward deployed engineering (FDE) engagements across several industries, with fictional companies and callers. Transcript excerpts use a few light conventions from CA transcription: numbers in parentheses are silences timed in seconds, aligned square brackets mark where caller and AI agent talk overlaps, and double parentheses hold transcriber comments.[5] Press play on any example to hear it rendered in voices available via ElevenLabs.

1. Managing perceived latency

Time is a design material in a voice interface. Timing carries meaning, silence is never neutral, and latency is arguably the most important usability challenge to design for, and around. In human conversation, we notice and often respond to delays on the order of milliseconds. Transitions between turns at talk typically occur with minimal gaps, often around 100–300ms, or even slight overlaps. Longer response times get noticed and interpreted; a gap means something, and interlocutors immediately start working out what (Sacks, Schegloff, & Jefferson, 1974; Kendrick & Torreira, 2015). Traditional cascaded voice AI architectures, which chain speech-to-text (STT), a large language model (LLM), and TTS, can easily add 1,000ms or more of latency.[6] The human expectation of near-zero-latency transitions is hard to meet, and doesn’t have to be. Instead, manage the silence. An unannounced silence beyond 500–1,000ms makes the agent appear unresponsive and disconnected.

You can mitigate this by designing for perceived latency using interactional feedback cues:

  • Verbal fillers: Short phrases (“Let me check that…”) that acknowledge the request before backend processing and tell callers that they are still connected.
  • Nonverbal audio: Short, subtle sounds streamed during processing gaps to reassure the caller that the system is active. These pair well with verbal fillers.[7]

Listen to example scenario 1

A WISMO (“Where is my order?”) agent, on an order status lookup with a 4-second backend delay.

Order status lookup, 4-second backend delay
Poor experience (unmanaged silence)
0:00 / 0:10
Caller: Can you check when my order is arriving tomorrow? (4.2) ((silence while lookup runs)) Caller: Hello? Are you [still there? Agent: [Your order is out for— Caller: Sorry, say that again?
Improved experience (managed perceived latency)
0:00 / 0:11
Caller: Can you check when my order is arriving tomorrow? (0.2) Agent: Let me track that. (3.8) ((beepboop audio cue plays)) Agent: Okay, it’s out for delivery tomorrow between 9:00 AM and noon.

For delays that outlast the acknowledgment and the audio cue, add incremental progress updates (“Still looking…”). For lookups in the tens of seconds, announce the duration, if possible.

Latency fillers are a micro-manifestation of persona. Keep their phrasing consistent with the agent’s personality and its prompt instructions. A caller shouldn’t hear one voice say “One moment, please” and the same voice, a turn later, say “Hang tight.”

Encoding it in Agentforce

In the voice channel configuration, lower beepboop_config from its 6,000ms default to about 500ms, so the filler audio reads as ambient “working” rather than as a sign something has gone wrong (see Choosing a voice above).

Agent Script: Latency: acknowledge before the lookup

subagent parcel_tracking:
  reasoning:
    instructions:
      | When a request needs a lookup, acknowledge it first: "Sure, let me track that." Never leave the caller in an unannounced silence.
      | When you deliver the result, open with a resumption marker such as "Okay" or "Alright".    
  actions:
      track: @actions.track_parcel
  actions:
    track_parcel:
      target: "<apex://ParcelTracking>"
      include_in_progress_indicator: True    # without this, the message never plays
      progress_indicator_message: "Let me track that."

The instruction asks the agent to acknowledge; progress_indicator_message guarantees it. The platform speaks that line the moment the action is dispatched, so it lands even when the reasoning engine decides to say nothing. include_in_progress_indicator defaults to False, so a progress_indicator_message on its own never plays. The result is the unmanaged silence in the transcript above. Both fields belong on the action definition, not on the binding in reasoning.actions. Phrase the message as talk, not as a status caption: “Let me track that,” not “Retrieving order data.”


2. Designing for turn-taking

Turn-taking (“who should talk next and when should they do so?”, Schegloff, 2007: xiv) is the structural framework of conversation, and it is precisely organized. Interlocutors project a turn’s upcoming completion from its grammar, prosody, and context instead of waiting for it to end (Sacks, Schegloff, & Jefferson, 1974). That projection is how human talk achieves the transition timings discussed above, and the pattern holds across languages (Stivers et al., 2009). The moment when a turn could be complete, and the listener may take over, is called a Transition Relevance Place (TRP).

The consequences for designing clean turn-taking are:

  • For agent turns: Project clear completion points. Use complete grammatical structures and falling intonation for declarative statements, and end on the exact item requiring a caller response. Trailing off or asking a question followed by further talk deprives the caller of a clear TRP. The result is overlap, or dead air that the caller doesn’t realize is now theirs to fill.
  • For caller turns: Recognize that “early” talk is normal. Callers often speak early when they anticipate where a turn is going. Treat barge-in, the design gloss for interruption and other kinds of overlap, as a sign of caller engagement rather than a system failure. When interrupted, the agent has to immediately yield the floor. Don’t talk over the caller or awkwardly restart a script (e.g., “As I was saying…”, repeating verbatim).

Listen to example scenario 2

Flight options delivery: an airline’s agent begins listing options. The caller, having projected where the list is going, short-circuits it:

Flight options delivery, caller short-circuits the list
Poor experience (ignoring overlap/barge-in)
0:00 / 0:07
Agent: I found three flights. The first departs at seven fifteen with [one stop in— Caller: [Just give me the cheapest one. (0.5) Agent: The second departs at ten twenty with…
Improved experience (yielding the floor & adapting)
0:00 / 0:16
Agent: I found three flights. The first departs at seven fifteen with [one stop in— Caller: [Just give me the cheapest one. (0.5) Agent: The cheapest is the ten twenty departure at two hundred forty dollars. Would you like to book that, or hear the others? Caller: That one. (0.6) Agent: Alright, the ten twenty flight for two hundred forty dollars…

Encoding it in Agentforce

In the voice channel, interruption detection is the platform’s job, so you don’t script it. The platform decides when the caller has taken the floor mid-utterance, halts playback, and hands the new turn to the agent for another round of reasoning. Barge-in is on by default in Agentforce Voice and has no sensitivity setting to tune. endpointing_config governs turn-final silence, meaning how long a pause ends the caller’s turn. The platform default is already tuned for this, and lowering it clips callers who pause mid-thought, so it’s best left alone. For this principle, then, the voice channel offers almost nothing to set.

In Agent Script, you can’t change whether or when an interruption registers, but you can control what the agent does once the platform surfaces one. The instructions below encode three behaviors: yield the floor, treat the barge-in as the new current turn, and don’t resume the interrupted utterance. Agent Script also lets you shape turns for interruptibility in the first place by keeping them short enough that a caller who jumps in isn’t cutting off something they needed to hear.

Agent Script: Turn-taking: yield on interruption

subagent flight_booking:
  reasoning:
    instructions:
      | If the caller speaks while you are mid-turn, yield immediately.
      | Treat the interruption as the current turn and answer it directly. Do not complete or resume your previous utterance.
      | Keep each turn to one or two short sentences, ending on the precise element the caller needs to address.

3. Composing turns: Information structure

How you build a turn matters as much as where you place it. CA treats every turn as analyzable for its position and its composition, and composition in a voice interface is governed above all by end focus: in spoken English, new and important information gravitates toward the end of the clause (see Cohen, Giangola, & Balogh, 2004). Because the turn-final element sits right before the TRP, it’s freshest in the caller’s working memory, and it most strongly sets the terms of that response. When an agent packs multiple questions into a single turn, callers tend to answer the last one. That isn’t inattention. The turn-final question is the one still making a claim on them when the floor opens.

With end focus in mind, here are some design recommendations for turn composition:

  • Place the question at the very end of the turn. Don’t follow a question with additional information.
  • Frame context first, news second: Put the context before the new or focal element (e.g., “Regarding your Thursday appointment: it’s been moved to 3 p.m.”).
  • Put caveats before the item they qualify: Stating a price followed by “nonrefundable” feels dishonest. Stating “This nonrefundable option is $100” frames the trade-off correctly.
  • Let information structure outrank style rules: Word order should serve information structure even against style-guide conventions like “no passive voice.” Use passive voice if it places key information at the end (e.g., “Your claim will be handled by our specialist team” properly emphasizes the team).
  • One action per turn: Deliver one piece of information, confirm one understanding, or ask one question. This is the safe default, especially for first-time callers who don’t yet know the shape of the process. Stepping through it one item at a time spares their working memory. Adapt it to the audience, though. Power users who run the same flow many times a day can anticipate the sequence and may prefer to volunteer several details at once, and some sequences are so familiar (name and date of birth in one breath) that most callers do the same. Design for that by capturing volunteered slots into variables rather than re-eliciting them (see Sparing working memory), so accepting more per turn never means asking twice.

Listen to example scenario 3

A healthcare booking agent, rescheduling an appointment:

Healthcare booking, rescheduling an appointment
Poor experience (question placed early)
0:00 / 0:04
Agent: Would you like to keep the ten o’clock with Dr. Popli? Be[cause I should mention— Caller: [Uh…when else is she free?
Improved experience (end focus applied)
0:00 / 0:09
Agent: Dr. Popli has an opening Thursday afternoon, and changing today incurs no fee. Would you like to keep ten o’clock, or switch to Thursday? Caller: Let’s switch to Thursday.

Encoding it in Agentforce

Turn composition depends almost entirely on the wording of prose prompt instructions, which should read like design rules for turns. Conditional logic places situational content in the right position.

Agent Script: Turn composition: end focus as a rule

variables:
  change_fee: mutable number = 0
    description: "Fee for changing this appointment today, in dollars."

subagent reschedule_appointment:
  reasoning:
    instructions: ->
      | Perform exactly one action per turn: deliver one piece of information, confirm one detail, or ask one question.
      | Place the item requiring a response at the very end of the turn. Never place text after a question.
     | State conditions, caveats, and fees BEFORE the item they apply to.
      if @variables.change_fee == 0:
        | Say that changing today carries no fee, then ask the caller to choose.

4. Sparing working memory

Speech is heard once and then gone, with no scrolling back.[8] That ephemeral or nonpersistent quality is the central constraint on how much a voice turn can carry. The auditory channel is transient and sequential, and everything the caller must act on has to survive in working memory until it’s their turn to speak. Voice design that respects this constraint converges, reassuringly, with how talk-in-interaction is organized.

An agent that packs three requests into a single turn opens three question-answer pairs (“adjacency pairs”; Schegloff, 2007) at once and asks the caller to service all of them from memory. Predictably, the caller answers the turn-final one (end focus again), and the rest tends to go missing. The agent then re-asks, and the interaction pays for the same information twice. That costs progressivity, the default expectation that talk keeps moving forward. Gathering information one question at a time reproduces the stepwise, incremental structure through which institutional encounters (e.g., medical intake, service counters) collect information anyway. If a caller volunteers several details at once, capture all of them into variables without forcing the caller to repeat them.

To offload working memory:

  • Cap lists at three items: Conversational lists recurrently follow a three-part structure (Jefferson, 1990). Avoid reading more than three options at once.
  • Deliver progressively: When the details would outrun working memory, give headlines first and the details on request (“You have four appointments next week. Want to hear the times?”). Below that threshold, just say it: “You have appointments tomorrow at ten and two” beats making the caller ask.
  • Choose between implicit and explicit confirmation:
    • Implicit confirmation echoes received information inside the next turn (“14 Rosemary Lane. Is morning or afternoon better?”). This validates a previous message with zero extra turns. The caller speaks up only if it’s wrong (“Actually, it’s 40 Rosemary Lane”), and the stored value is updated in place. Each detail is confirmed and saved as it arrives, with no “Is that right?” on every turn.
    • Reserve explicit confirmation (“Should I book this now?”) for hard-to-reverse or high-consequence (transactional) actions.

Listen to example scenario 4

Service visit scheduling.

Service visit scheduling
Poor experience (form-filling anti pattern)
0:00 / 0:08
Agent: I’ll need your street address, whether someone over eighteen will be home, and whether you prefer morning or afternoon. (0.3) Caller: Uh…afternoon?
Improved experience (sequenced, implicit + final explicit summary confirmation)
0:00 / 0:24
Agent: I can get that rescheduled. What’s your street address? Caller: 14 Rosemary Lane. Agent: 14 Rosemary Lane. Is morning or afternoon better? Caller: Afternoon. And the gate code is 4521. Agent: Afternoon, gate code 4521. Last thing: will someone over eighteen be home? Caller: Yes. Agent: Perfect. I have an installation at 14 Rosemary Lane tomorrow afternoon, gate code 4521. Should I book this now?

Encoding it in Agentforce

Slot-by-slot gathering needs state that outlives the turn, so it belongs in variables rather than in the model’s context. @utils.setVariables encodes this. It takes a whole set of slots at once, with ... meaning “the agent fills this from what was just said,” so a caller who volunteers “Afternoon, and the gate code is 4521” gets both stored in one pass and is never asked twice. The instruction still enforces one question per turn; the variables just make sure nothing volunteered early is lost. Capture happens turn by turn as each detail arrives, so the closing summary reads back stored values rather than re-eliciting them.

Agent Script: Working memory: slot-by-slot capture

variables:
  street_address: mutable string = ""
    description: "The service street address."
  time_preference: mutable string = ""
    description: "Morning or afternoon."
  gate_code: mutable string = ""
    description: "Gate or entry code."
  adult_present: mutable boolean
    description: "Whether an adult over 18 will be present. Unset means unasked."

subagent book_installation:
  reasoning:
    instructions: ->
      | Ask for exactly one missing detail per turn, echoing the previous answer inside the next question ("14 Rosemary Lane. Morning or afternoon?").
      if @variables.street_address != "" and @variables.time_preference != "" and
@variables.adult_present is not None:
        | Summarize {!@variables.street_address}, {!@variables.time_preference} gate code {!@variables.gate_code} in one sentence, then ask for explicit confirmation.
    actions:
      capture_details: @utils.setVariables
        description: "Store every detail the caller gives, even several volunteered in one turn. Never re-ask for something already stored correctly."
        with street_address = ...
        with time_preference = ...
        with gate_code = ...
        with adult_present = ...

5. Repairing by degrees

In human interaction, repair means fixing misunderstandings, mishearings, or misstatements (Schegloff, Jefferson, & Sacks, 1977). It’s a standard structural feature of conversation rather than an error state, and its basic organization is among the strongest cross-linguistic universals we have (Dingemanse & Enfield, 2024). Repair has its own preference structure: speakers prefer to repair their own talk, and when the recipient must step in (“other-initiated repair”), the initiations are graded from weak to strong.

The design translation is a rule worth enforcing mechanically. Use the weakest repair initiator that can locate the trouble, and escalate step by step:

  1. Open-class initiators (“Sorry?”, “Pardon?”): Use these when nothing recognizable was heard. They force the caller to repeat the entire turn.
  2. Category-specific initiators (“Which account, checking or savings?”): These locate the trouble source and ask for targeted clarification.
  3. Candidate understandings (“Did you mean cancel the order placed yesterday?”): These offer a specific guess for the caller to verify.

For names and long numbers, where the trouble is acoustic rather than semantic, switch encoding: spelling, digit-by-digit, or another modality (see next section).

A few rules apply at every level of repair:

  • Never repeat a failed prompt verbatim: If a prompt failed, repeating it implies the caller is at fault or hard of hearing.
  • Never blame the caller; attribute trouble to the system: “I didn’t catch that,” not “Your response was not recognized.”
  • Treat silence as its own trouble type: Nonresponse warrants a polite pursuit (Pomerantz, 1984). First a gentle re-ask, then a reformulated choice, then an alternative channel.
  • Count global failures across the call: A caller who has hit trouble on three different questions is telling you the voice channel is not working for them today. Proactively offer a transfer to a human rep.

Listen to example scenario 5

Bank account name lookup over a noisy connection.

Bank account name lookup, noisy connection
Poor experience (verbatim loop)
0:00 / 0:11
Agent: What’s the name on the account? Caller: ((Name, against traffic noise)) (0.7) Agent: I’m sorry, I didn’t understand. What’s the name on the account? Caller: ((Repeats, louder)) (0.6) Agent: I’m sorry, I didn’t understand. What’s the name on the account? Caller: Forget it.
Improved experience (graduated repair)
0:00 / 0:15
Agent: What’s the name on the account? Caller: ((Name, against traffic noise)) (0.7) Agent: Sorry? Caller: (Repeats name, against traffic noise)) (0.6) Agent: Sorry, I didn’t catch that. Could you spell your last name for me? Caller: N-G-U-Y-E-N. (0.8) Agent: N-G-U-Y-E-N, Nguyen. Got it. And your first name?

Encoding it in Agentforce

The four-step ladder (three initiators, then a switch in encoding) is instruction wording. The repair budget isn’t a judgment call, so it goes somewhere the reasoning engine can’t reach. after_reasoning runs after every turn, unconditionally. Once the counter hits three, the call moves to a human rep whether or not the agent thinks it’s making progress. Escalation is a transition to a subagent that binds @utils.escalate. The counter can’t route anywhere on its own, so the target has to exist.

Agent Script: Repair: graduated ladder with a hard escalation budget

connection messaging:
  outbound_route_type: "OmniChannelFlow"
  outbound_route_name: "flow://EscalationFlow"

variables:
  repair_attempts: mutable number = 0
    description: "Counts turns where the caller could not be understood."

subagent account_lookup:
  reasoning:
    instructions:
      | If nothing recognizable was heard, ask once: "Sorry?"
      | If one element is unclear, name it and offer specific choices.
      | If you have a strong candidate, verify it ("Did you mean...?").
      | For names and long numbers, ask for spelling or digit-by-digit.
      | Never blame the caller. Never repeat a failed prompt word-for-word.
    actions:
      count_trouble: @utils.setVariables
        description: "Add one to repair_attempts on any turn you could not make out. Do not call it on any other turn."
        with repair_attempts = ...
  after_reasoning:
    if @variables.repair_attempts >= 3:
      transition to @subagent.handover

subagent handover:
  reasoning:
    actions:
      to_human: @utils.escalate

The escalation action hands the call to a human rep through your Omni-Channel routing, ideally with the transcript and collected variables attached, so the caller doesn’t repeat themselves. For acoustic trouble like the noisy line above, where the problem is hearing rather than meaning, one field helps with the vocabulary you can predict. Speech recognition assigns low prior probabilities to proper names and domain terms, and those make up much of what an enterprise agent has to hear correctly. inbound_keywords in modality voice: biases recognition toward a list you supply. Fill it with the vocabulary you can predict or know from testing (product names, plan tiers, branch locations, etc.). Caller surnames aren’t enumerable. For those, the repair ladder above is a solution. In the same block, inbound_filler_words_detection drops filler tokens like um and uh when set to True (often useful for telephony, but worth experimenting with). Leave it off where a hesitation is a meaningful signal.

Agent Script: Key-term biasing for speech recognition

modality voice:
  inbound_keywords:
    keywords:
      - "Coastline Checking"
      - "AutoSweep"
      - "Harborview Savings"
  inbound_filler_words_detection: True

6. Going multimodal where it helps

Voice excels at hands-free, eyes-free, turn-by-turn exchanges with modest information density. It’s weak when dealing with comparative evaluations (comparing three or more options held in mind at once) or character-exact data entry (e.g., alphanumeric tracking codes and order IDs, complex URLs, email addresses).

Five signals suggest a task may be a poor fit for voice:

  • The caller must hold three or more options in mind to compare them.
  • The input must be character-exact.
  • The output must persist beyond the call (e.g., terms to re-read, a reference number).
  • The turn must gather several exact fields at once, or capture a document or image.
  • The same slot keeps failing repair (Principle 5’s counter).

Any one of them, or any combination, is a reason to consider switching.

Instead of forcing these tasks through the ear, switch modality. Hand them to a screen (an email, an SMS link to a pre-filled form, or a comparison view in the app), or use the keypad via dual-tone multi-frequency (DTMF) for character-exact numeric entry, which sidesteps speech recognition entirely. Callers already understand keypad input from decades of interactive voice response (IVR) systems. Reusing that convention for exact numbers is a good example of using the cleanest channel for the job, and is not a step backward to old phone trees.

When handing off to another channel:

  • Get explicit confirmation first: Ask before sending (“Want me to text a link?”).
  • Make the context travel: Pre-fill every field already collected. Context loss is one of the most common cross-channel failures in field engagements.
  • Keep the door back to voice open: Stay on the line while they complete entry.

Listen to example scenario 6A

A patient pre-check-in that needs several character-exact fields and an insurance card photo:

Patient pre-check-in (several exact fields + insurance card photo)
Poor experience (forcing many exact fields through voice)
0:00 / 0:25
Agent: Before your visit I need your date of birth, your insurance member ID, the group number, and your pharmacy. (1.2) Caller: It’s March fourteenth 1985, and the member ID is R, X, Z, zero… or is that an O? Agent: Sorry, was that R as in Romeo? Caller: This is going to take a while, isn’t it? (0.9) Agent: Sorry, I didn’t catch that. I need your date of birth, your insurance member ID, the group number, and your pharmacy.
Improved experience (multimodal handoff)
0:00 / 0:19
Agent: There are a few details to check you in, and a photo of your insurance card. I can text a secure link so you can type them in and take a picture of the card. Want me to send it? Caller: Yes. (3.4) ((beepboop audio cue plays)) Agent: Sent. I’ll stay on the line while you finish, let me know when the card’s attached.

Several exact fields plus an attachment clearly require a multimodal handoff. A single field is a subtler call. Take an email address: It’s usually better to look it up than to have the caller spell it out. If you can cleanly capture an identifier the caller says or dials, such as a phone number, account number, or order number, you can retrieve the address on file and confirm it in a single turn.[9] A phonetic-alphabet check (“Was that M as in Mike?”) can still rescue a stray character, since near-homophones like M and N are hard for speech recognition (see Principle 5). Reserve spelling, or a texted link, for when there’s no identifier to cross-reference against.

Listen to example scenario 6B

Character-exact numeric entry, an order number via keypad:

Collecting an order number by keypad (DTMF)
Keypad entry (DTMF)
0:00 / 0:15
Agent: I can look that up. You can speak or use your keypad: what’s your eight-digit order number? (4.9) ((caller keys eight digits)) (2.0) ((beepboop audio cue plays)) Agent: Got it, order ending in six six three one. It’s out for delivery tomorrow.

Keyed digits arrive as data, not audio, so there is nothing for speech recognition to mishear. The read-back is partial (the last four digits), and the nonverbal audio cue manages the latency. For sensitive numbers (e.g., a card, a PIN), the keypad is also a privacy device; what’s keyed isn’t overheard.

Encoding it in Agentforce

The handoff is an action, plus the instructions governing when to offer it, plus two guards on the consent. target: is required on every action definition. Here it’s a Flow that sends the SMS and builds the pre-filled link, so the URI depends on your org’s messaging setup.

Agent Script: Multimodal: consented handoff to SMS

variables:
  handoff_accepted: mutable boolean = False
    description: "Whether the caller agreed to be texted a link."

subagent order_wrap_up:
  reasoning:
    instructions:
      | When a task needs several exact fields, or a document or photo, offer to text a secure link instead.
      | Ask before sending, say what you are sending, confirm when it is gone.
     | Stay on the line and offer help while the caller uses the link.
    actions:
      accept_handoff: @utils.setVariables
        description: "Set to True when the caller agrees to be texted a link."
        with handoff_accepted = True
      send_link: @actions.send_prefilled_link
        available when @variables.handoff_accepted == True
  actions:
    send_prefilled_link:
      target: "flow://Send_Prefilled_Form_Link"
      require_user_confirmation: True

The keypad path is simpler because there is nothing DTMF-specific to encode. Keyed digits reach the reasoning engine as plain text, identical to a transcribed spoken turn, so the same variable capture stores the number. The design centers on the wording of prose instructions: whether you offer the keypad, when you nudge toward it, and where in the turn it sits.

Agent Script: Keypad entry: capture an order number

variables:
  order_number: mutable string = ""
    description: "The caller's eight-digit order number, spoken or keyed."

subagent order_lookup:
  reasoning:
    instructions:
      | When you ask for the order number, offer both channels, options first, question last: "You can speak or use your keypad: what's your eight-digit order number?"
      | Treat keyed digits exactly like spoken ones. Never ask which way the number arrived, and never re-ask for a number you already have.
     | For sensitive numbers such as card numbers, suggest the keypad.
    actions:
      capture_order: @utils.setVariables
        description: "Store the eight-digit order number, whether spoken or keyed."
        with order_number = ...
      track: @actions.track_order

Putting it into practice

Design the voice channel configuration in the same pass as the instruction wording. Latency messages, voice selection, pronunciations, and escalation wiring should stay consistent with the script, and several of them limit what the script can fix.

Then test with audio, not text. Typing test utterances into a preview verifies the reasoning, but it misses barge-in behavior, background noise, accents, and timing. Those are the areas where voice deployments tend to struggle. Call the agent from a car, from a kitchen with the range hood running, from a room with a television on. Hand the phone to someone who has never seen the script and ask them to book something. Listen for the places where you would have interrupted and the agent kept talking, and for the silences you didn’t design. Finally, add scenario suites to Agentforce Testing Center so regressions surface when you edit subagents.

Design is now the bottleneck in voice AI, and design is encodable: conversational behavior as concrete rules in prompt instructions, sequential guarantees in variables and the deterministic blocks around the reasoning loop, and sound, timing, and wiring in the voice channel configuration. These principles focus on the timing and shape of turns. How turns sound, including voice selection, pronunciation, and accent, belongs to agent persona design and deserves the same care.

Callers can’t unlearn how conversation works. We conduct much of our lives in conversation, from early childhood on, and we bring to every conversation a web of shared conventions and expectations. Some are universal (Stivers et al., 2009), and others are specific to local language communities. Most of these expectations are “seen but unnoticed” (Garfinkel, 1967). When they hold, interaction feels natural; when they don’t, it feels off. Design with the rules of conversation, encode them into Agent Script, and you come closer to agents that work the way conversation does.

Everything You Ever Wanted to Know About Agentforce Voice (But Were Afraid to Ask)

Take these principles from design to deployment. The guide covers Agentforce Voice architecture, latency benchmarks, and escalation design, plus five use cases live in production today.


Notes

  1. How Voice AI Is Reshaping the $135 Billion Call Center Industry
  2. This guide focuses on voice-only interfaces for telephony applications, though many of the same design principles, with appropriate refinement, carry over to multimodal interface design.
  3. While the guide (Lucy, 2026) focuses primarily on agent persona design for text-based chat agents, it offers a rigorous framework that can be adapted and refined for voice.
  4. Detailed considerations around multimodal interface design, caller identification and authentication flows, multilingual and code-switching behavior, and the design of call openings and closings are deliberately out of scope.
  5. These are simplified from the Jeffersonian transcription system, the standard in conversation analysis. The handful used here is all a reader of this guide needs.
  6. More recent native speech-to-speech models have shown much lower latency, with sub-200ms response times.
  7. The voice AI industry sometimes offers “background noises” like keyboard tapping as latency management devices. Simulating human presence this way, or suggesting that it be simulated, is ethically problematic: it misleads callers about whether a human is present (and edges into uncanny valley territory), which reinforces the need for AI self-disclosure as part of responsible AI use.
  8. Voice interfaces are the purest expression of what current AIUX design sometimes terms an “ephemeral interface.”
  9. An ANI (Automatic Number Identification) lookup is often a good solution for this look-it-up move: the calling number arrives with the call, before the first turn, so a matching contact record can be pre-fetched and its details offered for confirmation. The caller confirms what’s already known, instead of providing it again.

References

Cohen, M. H., Giangola, J. P., & Balogh, J. (2004). Voice user interface design. Addison-Wesley.

Dingemanse, M., & Enfield, N. J. (2024). Interactive repair and the foundations of language. Trends in Cognitive Sciences, 28(1), 30–42.

Garfinkel, H. (1967). Studies in ethnomethodology. Prentice-Hall.

Jefferson, G. (1990). List construction as a task and resource. In G. Psathas (Ed.), Interaction competence (pp. 63–92). University Press of America.

Kendrick, K. H., & Torreira, F. (2015). The timing and construction of preference: A quantitative study. Discourse Processes, 52(4), 255–289.

Lucy, N. (2026). Agent personas for Agentforce: Designing agent personality and encoding it into Agentforce. Salesforce.

Pomerantz, A. (1984). Pursuing a response. In J. M. Atkinson & J. Heritage (Eds.), Structures of social action: Studies in conversation analysis (pp. 152–163). Cambridge University Press.

Sacks, H., Schegloff, E. A., & Jefferson, G. (1974). A simplest systematics for the organization of turn-taking for conversation. Language, 50(4), 696–735.

Schegloff, E. A. (2007). Sequence organization in interaction: A primer in conversation analysis. Cambridge University Press.

Schegloff, E. A., Jefferson, G., & Sacks, H. (1977). The preference for self-correction in the organization of repair in conversation. Language, 53(2), 361–382.

Stivers, T., Enfield, N. J., Brown, P., Englert, C., Hayashi, M., Heinemann, T., Hoymann, G., Rossano, F., de Ruiter, J. P., Yoon, K.-E., & Levinson, S. C. (2009). Universals and cultural variation in turn-taking in conversation. Proceedings of the National Academy of Sciences, 106(26), 10587–10592.

Stokoe, E., Albert, S., Buschmeier, H., & Stommel, W. (2024). Conversation analysis and conversational technologies: Finding the common ground between academia and industry. Discourse & Communication, 18(6), 837–847.

Further reading

Jefferson, G. (1986). Notes on “latency” in overlap onset. Human Studies, 9(2–3), 153–183.

Moore, R. J., & Arar, R. (2019). Conversational UX design: A practitioner’s guide to the natural conversation framework. ACM Books.

Pearl, C. (2016). Designing voice user interfaces: Principles of conversational experiences. O’Reilly Media.

Acknowledgments

Chris Rolfe — for reviewing and testing the Agent Script sample code, and for bringing important client engagement learnings to the draft.

Claude Sutterlin — for reviewing an early draft, sharpening the code examples, and pushing me to publish.

Gauthier Muguerza — merci for your Agentforce Voice product expertise that guided the encoding model.

Nathan Lucy — a big thank-you for thoughtful feedback on the earliest of drafts, and for encouraging me to write this up.

Bunly Lay, Alan Khan, Brad Shapiro, Mattias Sjövall, Susane Antunes, and the broader FDE team (too many to name, but you know who you are) — your field experience directly shaped the thinking behind this paper.

Agentic Experience Specialists — thank you for urging me to write this down, giving editorial feedback, and championing voice UX with our clients.

Get the latest articles in your inbox.