Chrome 150 On-Device Speech: Choose the Smallest Quality Mode That Fits the Job

Choose Chrome 150 command, dictation or conversation speech quality, check local model availability and design a disclosed fallback path.

Web accessibility lab comparing command, dictation, and multi-speaker conversation speech tasks
Choose the speech job before the local model.

Command, dictation, and conversation are not a simple low-to-high slider. They describe different speech jobs, local resource needs, and acceptance tests.

A voice button that recognizes “next” does not need to behave like a meeting transcript. A notes field that captures one speaker does not need to separate a roundtable. Chrome 150’s on-device Web Speech quality levels make that distinction explicit: command, dictation, and conversation.

The winning mode is not the largest one a device can download. It is the smallest mode that passes the product’s actual speech job—with a disclosed fallback when the local capability is unavailable.

Three physical language-pack models for command, dictation, and conversation speech recognition with a separate fallback card
The workload grows from short control phrases to continuous single-speaker text and then multi-speaker conversation. Fallback is a separate product decision, not an invisible automatic upgrade. Original Neyrotex illustration.

Round one: command wins when vocabulary is bounded

Use command quality for short control phrases: “play,” “pause,” “next slide,” or a constrained set of application actions. The acceptance test should emphasize false activations, response time, and behavior under ordinary room noise. Long-form punctuation is not the job, so scoring it only inflates the test.

Round two: dictation wins when one person owns the text

Use dictation for continuous speech from a single speaker: notes, messages, form fields, or document input. Test sentence boundaries, editing flow, proper nouns relevant to the domain, and whether the user can correct errors without abandoning voice. A command model that recognizes keywords is not a safe substitute for faithful text capture.

Round three: conversation wins when turns matter

Use conversation quality when several speakers or conversational context are part of the job. That raises the resource and product stakes. The browser’s ability to install or expose a local capability does not settle consent, speaker labeling, retention, or what happens when diarization is uncertain.

Choose by product contract, not by the impressive label
Job Start with Acceptance evidence Honest fallback
Voice controls Command Known intents, false-trigger rate, response time Visible touch or keyboard controls
Single-speaker text Dictation Editable transcript, punctuation, domain terms Manual entry or disclosed remote recognition
Multi-speaker session Conversation Turn boundaries, speaker uncertainty, long-session behavior Single-speaker mode or no recording
Availability is necessary; a job-specific acceptance test is what makes the mode releasable.

Check capability before promising local speech

The current API surface is experimental and availability can depend on browser version, language support, permissions, and a local model download. Query the documented availability state before starting recognition. Treat “downloadable” differently from “available”: the first may require time, storage, and a user-visible action.

Do not silently send audio to a server when local recognition is missing. A fallback that changes the processing boundary needs a clear disclosure and a user choice. If the product cannot offer that choice, keep the keyboard path primary.

Owner handoff before shipping

  • Product: names the speech job and fallback.
  • Engineering: checks language-pack availability and handles download, denial, and interruption.
  • Design: keeps a visible non-voice control and communicates processing state.
  • Privacy: approves the boundary for audio and transcript data.
  • QA: runs the mode-specific acceptance set plus an unavailable-pack negative control.

Official sources