by datastudy.nl

Friday, July 24, 2026

AI

Claude voice mode expands to Opus and Sonnet models

Claude voice mode is Anthropic's spoken AI interface. It now supports Opus and Sonnet for deep reasoning, plus Gmail, Slack, and nine new languages.

Slope chart showing Claude voice mode expansion from 2025 to 2026. Models with voice support rose from 1 to 3, languages from 1 to 10, and app integrations from 0 to 3. Anthropic added Opus, Sonnet, Gmail, Slack, Canva, and nine new languages.
Claude voice mode capability expansion, 2025 to 2026. Models with voice support went from 1 (Haiku) to 3 (Haiku, Sonnet, Opus). Languages went from 1 (English) to 10. App integrations went from 0 to 3 (Gmail, Slack, Canva). Source: The Verge. Data Today benchmark.

Talking to your AI assistant used to mean choosing between speed and intelligence. You could get quick answers from a lightweight model, or you could switch to text and type out a complex problem for the heavy lifter. That tradeoff just got smaller.

Claude voice mode now runs on Opus and Sonnet, turning spoken AI into a workflow tool.

Anthropic has expanded Claude voice mode beyond its fastest model. The spoken interface launched in 2025 on Claude Haiku, Anthropic's lightweight tier. It now supports Opus and Sonnet, the company's deeper reasoning models. The expansion also adds app integrations with Gmail, Slack, and Canva, and stabilizes voice support in nine new languages: French, German, Spanish, Hindi, Indonesian, Italian, Japanese, Korean, and Portuguese. For builders, the move matters because voice has become a full interface for complex AI workflows.

What did Anthropic actually change with Claude voice mode?

Three things shifted at once: model access, app reach, and language coverage.

First, the model tier. Voice mode was locked to Haiku, which Anthropic positioned as the fast, lightweight option. The company said in its announcement that people immediately pushed voice mode beyond casual queries, using it to work through real business problems. Haiku, in Anthropic's own words, "kept conversations quick, but not always deep." Sonnet and Opus, which Anthropic describes as "designed for hard problem-solving," can deliver more complex responses and take agentic actions: turning a conversation into a one-page pitch, or reshuffling calendar appointments when a train runs late.

Second, app integration. Voice mode can now reach into Gmail, Slack, and Canva. A spoken conversation can trigger actions in tools a builder already uses, rather than just returning text in a chat window. The Verge reports that Anthropic is positioning these integrations as part of a broader push toward voice-driven agentic workflows.

Third, language coverage. Nine languages move out of beta into full support, bringing the total to ten languages including English. The beta phase for non-English languages was limited; now French, German, Spanish, Hindi, Indonesian, Italian, Japanese, Korean, and Portuguese are all generally available.

Two more details matter for builders. Users can switch between text and voice mid-conversation without losing context. They can also change models mid-conversation, so a quick Haiku exchange that sparks an idea can seamlessly escalate to Opus for deeper exploration. That model-switching capability is unusual in the voice AI landscape. Most platforms lock you to one model per session.

Radar chart comparing Claude Haiku, Sonnet, and Opus across five dimensions. Haiku scores highest on response speed (5) and cost efficiency (5), lowest on reasoning depth (2) and agentic action (2). Opus scores highest on reasoning depth (5) and agentic action (5), lowest on cost efficiency (1) and speed (2). Sonnet scores mid-range on all dimensions (3 to 4). All three score 5 on language support.
Claude Haiku, Sonnet, and Opus compared across five dimensions relevant to voice mode. Haiku scores 5 on speed and cost efficiency, 2 on reasoning and agentic action. Opus scores 5 on reasoning and agentic action, 1 on cost efficiency, 2 on speed. Sonnet scores 3 to 4 across the board. All three score 5 on language support (10 languages). Values are illustrative, based on Anthropic's model descriptions as reported by The Verge. Data Today analysis.

The chart above shows how the three models compare across the dimensions that matter for voice interactions. Haiku leads on speed and cost efficiency but trails on reasoning depth and agentic capability. Opus is the opposite: deepest reasoning, highest cost. Sonnet sits in the middle on every axis.

Why does deeper voice reasoning change the calculus for builders?

With Opus and Sonnet in the mix, spoken interactions can now carry the weight of real analytical work.

Consider what a product manager could do. Start a voice conversation while walking to a meeting, ask Opus to analyze a competitor's pricing page, get a structured breakdown, then ask it to draft a Slack message to the team summarizing the findings. All spoken, all without opening a laptop. The app integrations with Gmail and Slack mean the output reaches the tools where work actually happens, rather than dying in a chat transcript.

The cost implications deserve attention. Opus sits at the top of Anthropic's pricing tier. Running it through voice, which tends to produce more back-and-forth turns than text, could rack up usage faster than a team expects. This is the same agentic cost problem that has bitten teams deploying coding agents: more capable models do more, but they also spend more. A team that gives every salesperson always-on voice access to Opus should model the token consumption carefully before rolling out broadly.

The mid-conversation model switch is the feature that tempers the cost concern. A team could default to Sonnet for most voice interactions, which sits in the middle of Anthropic's pricing, and reserve Opus escalation for moments that genuinely need it. That kind of tiered routing is already common in text-based AI deployments. Bringing it to voice is new.

For builders shipping voice-enabled products, the app integration angle matters most. If your product lives inside Slack or connects to Gmail, Anthropic's voice mode can now reach it. That changes the surface area for voice-driven automation. A user could verbally instruct Claude to pull data from a Slack channel, analyze it, and post a summary back. Whether that is empowering or alarming depends on how much you trust autonomous actions triggered by speech. The same agent security gap that has already hit half of enterprises applies: voice is just another entry point for an agent with tool access.

How does this stack up against competing voice AI products?

Anthropic is arriving later than OpenAI and Google to deep voice AI. Both competitors have offered voice features in their AI assistants, giving them more time to refine the experience, handle edge cases around accents and ambient noise, and build user habits. OpenAI's voice mode for ChatGPT has been available since 2024, and Google has integrated voice across its Gemini product line.

Anthropic's differentiator is the combination of model switching and app integration. The ability to start a voice conversation on Haiku, escalate to Opus when the problem gets hard, and have the output land in Slack or Gmail is a workflow that stands out. The mid-conversation model switch is a feature that most voice AI products do not offer in a single session.

The language expansion also narrows a gap. The specific set of languages Anthropic added, including Hindi, Indonesian, and Korean, targets markets where voice-first AI adoption may outpace text-first adoption. Mobile-first users in India and Southeast Asia are more likely to talk to an AI than type to one. Anthropic's decision to stabilize support for these languages, rather than keeping them in beta, signals a bet on those markets.

The competitive question is whether model switching and app integrations are enough to overcome the head start. ChatGPT's voice mode has had months to build user habits and gather conversational data that improves the experience. Anthropic is betting that developers and power users will value the ability to route between models more than they value polish.

What should you do with Claude voice mode now?

The practical moves depend on what you are building.

If you are building internal tools on top of Claude's API, test the voice interface as a first-class input method. The model switching capability means you can architect workflows that start cheap and escalate only when needed. Build routing logic that defaults to Sonnet for voice and escalates to Opus based on conversation complexity signals.

If you are evaluating voice AI for customer-facing products, the app integrations are the headline feature. A voice conversation that can write to Slack, draft a Gmail response, or update a Canva design is a different product from a voice chatbot that only returns text. Map the integration surface against your existing tool stack before committing.

If you are managing costs, model the token consumption of voice interactions carefully. Voice produces more turns per session than text because spoken conversations include filler, clarifications, and restarts. A voice session on Opus could easily consume two to three times the tokens of a text session covering the same ground. That estimate is directional, not measured, but the principle holds: voice is chattier, and chattier means pricier.

If you are hiring or building a team around conversational AI, the skill set is shifting. Voice UX design, accent and noise robustness, and latency optimization matter more than they do for text-only products. Anthropic putting its most capable models behind voice means the quality bar for voice AI experiences is rising across the board.

The voice tier is settling

Anthropic's move closes a gap that was starting to look conspicuous. Voice mode on Haiku only was a limited tool for a limited job. With Opus and Sonnet available, the question is now whether your workflows are ready for voice. The model switching and app integrations give builders a genuine new surface to design for. The cost question is real, but manageable with tiered routing. The competitive race is far from over, and Anthropic's bet on depth plus integration is a credible hand. The next twelve months will show whether developers build for it.

Sources

  • The Verge - Claude's voice mode is now available for Opus and Sonnet