How-to Guide

Voice Input Modes — choosing and building open-mic vs. push-to-talk

Design guidance for app builders deciding how a viewer's input reaches the agent. Three modes:

Mode Session class Mic permission How a turn starts
Open-mic (VAD) KalturaAvatarSession Prompted at connect* Viewer just speaks; server VAD cuts the turn
Push-to-talk KalturaAvatarSession (isTapToTalk) Prompted at connect* Viewer opens/closes a capture window (startTapToTalk()/endTapToTalk())
Chat (text-only) KalturaChatSession Never requested — the class never touches getUserMedia or WebRTC sendText('…') over plain HTTPS

* In the default micStartMode: 'immediate' the permission prompt runs alongside the connect handshake. connect() never waits for it and never fails because of it: a denied, missing or busy mic emits one warning (mic_permission_denied, mic_not_found, mic_in_use) and the session connects mic-less, with typed turns working. Or deferred: micStartMode: 'deferred' connects the avatar session with no mic at all. The app calls startMic() later from a real user click, so the permission prompt is gesture-anchored instead of firing on page load. Until then, typed turns work, but startTapToTalk() throws mic_not_started. See README.md.

The first two are voice-capture modes on the live avatar transport. Most of this doc is about choosing between them and building their UI.

Chat mode is a full text-only transport, with the same brain, same tools, and same thread as the voice modes. It's covered in README.md, including mid-conversation switching between avatar and chat via KalturaAgentSession (switching section below).

For the wire mechanics of startTapToTalk()/endTapToTalk(), see README.md. For the exact socket events, see Wire Protocol (tapToTalkStart/tapToTalkEnd, isTapToTalk).

Text input is not voice-mode-exclusive: an avatar session accepts typed turns too (session.speak(text) — see Client-Side Commands and README). Offering a text box alongside either voice mode is how the SDK satisfies EN 301 549 §6.2.1.2 (concurrent voice and text) — see README.md.

The one rule that overrides everything else here

Pick one VOICE mode per agent, at configuration time. Never offer both voice modes live in the same session. (Chat text-only input is exempt — typed turns don't touch the VAD/capture machinery this rule protects, so text can coexist with either voice mode, or replace voice entirely.)

This is not a UI-polish preference. It is a correctness requirement.

Don't rely on anything else to catch a mismatch for you. The SDK's own client-side gate (startTapToTalk() throws capability_disabled unless session.capabilities.tapToTalk) stops your UI from opening a tap-to-talk window on an agent configured for open-mic. Build your UI to match the agent's single configured mode. Never render a tap-to-talk control against an open-mic agent, or vice versa. Mixing them produces double-cut or missed turns (see Wire Protocol's tapToTalkStart/tapToTalkEnd row for the event pair involved).

Every push-to-talk/open-mic product draws the same line — one active capture mechanism, chosen once, not a live per-session toggle exposing both:

Product How it handles capture mode
Discord Single Input Mode setting — no per-session switch
Amazon Alexa / Ford SYNC Wake-word vs. PTT-button are alternate triggers for one active capture mechanism, not two concurrent ones
WhatsApp Hold-to-record voice notes only — no separate open-mic mode

Deciding which mode fits your app

Use open-mic (VAD) when… Use push-to-talk when… Use chat (text-only) when…
Viewers ask longer, exploratory questions (investor Q&A, tutoring, free-form conversation) Utterances are short, command-like bursts (a wake word, a walkie-talkie-style call) The viewer can't or won't speak (open office, quiet space, no mic) or won't grant mic permission
The environment is relatively quiet / single-speaker (a kiosk, a 1:1 demo) The environment is noisy or multi-speaker, and VAD would false-trigger on background talk Bandwidth is constrained — no video, no WebRTC, just HTTPS request/response
You want zero-friction "just speak naturally" — no button to find or learn Viewers need an explicit, deliberate boundary on when the mic is live (privacy-sensitive settings, shared/public devices) You need a zero-permission-prompt entry point; the viewer can switch up to the full avatar later without losing the thread

If your viewers will regularly speak in full sentences or ask multi-part questions, open-mic is the better default. Neither mode is inherently better. Press-and-hold can strain a viewer's hand on a long dictation, and VAD-based toggle can cut a speaker off mid-thought on long unstructured reasoning. Which complaint you'll hear from your own viewers tracks utterance length, not a fixed advantage of one mode over the other.

UX pattern: click-to-toggle, not press-and-hold

Prefer a single click/tap to open the capture window and a second click/tap to close it, over press-and-hold-to-record. Reasons, in order of weight:

If your product genuinely needs walkie-talkie-style short bursts (not this SDK's typical investor-deck or knowledge-avatar use case), a press-and-hold-with-slide-to-lock hybrid (WhatsApp's pattern) is the documented middle ground. Start from toggle, and only move to hold-based capture if you have evidence your utterances are consistently short.

Visual and non-visual state feedback

Give the viewer three redundant signals that the mic is live, not one:

  1. Icon/color change on the button itself: swap an idle-mic icon for a recording icon on a tap-to-talk button, not just an aria-pressed attribute change.
  2. An animated level indicator: a waveform, pulsing glow, or (simplest, and already available) this SDK's localMicLevel event ({level}, 0–1, ~50ms tick) driving a CSS custom property. A single static icon alone is not enough (NN/g's critique of Amazon Echo's single light ring as "a far cry from rich textual feedback").
  3. A live-region text or caption update (aria-live) confirming state changes ("Listening…" / "Sent") — necessary for screen-reader users who can't see the icon/waveform at all. Pair with an audio cue (a short start/stop tone) as an additional non-visual channel if your app's audio design allows it.

Safety: don't let a capture window hang open forever

A tap window that never closes (tab closed mid-recording, app crash, network drop) must not leave the server's tap-mode state stuck open. Layer these on top of startTapToTalk()/endTapToTalk():

Implementation checklist

  1. Decide the agent's mode at provisioning time — isTapToTalk is a fixed per-agent deployment choice (set wherever your app builds its intellect/session config), not something your UI code branches on live.
  2. Build the UI conditionally on session.capabilities.tapToTalk (derived from clientConfiguration.isTapToTalk). Render a tap-to-talk control only when true. Never render both a tap-to-talk control and an open-mic affordance for the same session.
  3. Wire click-to-toggle calling session.startTapToTalk()/session.endTapToTalk(), updating aria-pressed and an icon/level indicator on tapToTalkStarted/tapToTalkEnded — see the conditional-UI example in README.md.
  4. Add the silence/max-duration/abandonment safeguards above; the SDK does not impose them for you.
  5. Update any UI copy that currently assumes continuous listening (e.g. a text-input placeholder like "...or just speak naturally") — that copy is wrong for a tap-to-talk agent and should describe the button instead.

Switching between avatar and chat mid-conversation

KalturaAgentSession runs one conversation over either transport and switches between them with switchMode('avatar' | 'chat') — the thread, the request_vars context, and every onToolCall handler carry over automatically. One state machine:

From Call / event To Notes
idle connect() connecting → connected Once-only; a second connect() throws invalid_state
connected switchMode(other) switching → connected Emits transportChanged, then modeChanged {mode, threadContinuity}
connected switchMode(current) connected Idempotent no-op — nothing tears down
switching switch fails failed (reason: 'transport_failed') No rollback; buffered sends reject with the switch error
switching disconnect() closed The in-flight switchMode() rejects with invalid_state; one ended {reason:'disconnected'}, no failed
connected transport dies failed + ended forwarded Socket drop, server end
any disconnect() closed Idempotent; exactly one ended {reason:'disconnected'}

modeChanged.threadContinuity tells you which happened. true means the new transport was seeded with the live thread (show "conversation restored"). false means no turn had happened yet, so there was no thread to carry.

Two UX rules for the switch:

Doc What it adds
README.md The SDK API: startTapToTalk()/endTapToTalk(), the capability_disabled gate, tapToTalkActive/capabilities.tapToTalk
Wire Protocol The exact tapToTalkStart/tapToTalkEnd socket events and the finding behind the mixed-mode gate
Client-Side Commands A different silent client→page channel (tool calls), not voice input — useful contrast for what this doc is not about
README.md KalturaChatSession (text-only transport) and KalturaAgentSession (mode switching) — full API
Click to talk with Nova — she knows this whole SDK.
Nova AI assistant — knows this whole site

Reloading starts a fresh chat. “New conversation” does the same without leaving the drawer.