ElevenLabs and Google Gemini 3.8 reshape voice AI choices
As ElevenLabs' CEO lays out his pricing philosophy and stance on AI disclosure, Google's Gemini 3.8 TTS release sharpens the question of which voice AI platform actually suits your use case.

Two things happened in the same week that, taken together, tell you a lot about where the voice AI market is heading. ElevenLabs CEO Mati Staniszewski sat down at the Nrth conference in Toronto and spoke candidly about margins, IPO timing, and whether businesses owe their customers a disclosure when a bot picks up the phone. At almost the same moment, on 23 September, Google shipped Gemini 3.8 TTS and Gemini 3.8 Flash-Lite TTS, describing them as its most capable audio generation models to date.
Neither event alone is a revolution. But together they draw a sharper picture of a market where anyone choosing a voice AI tool, whether for a contact centre, an audiobook pipeline, or a gaming product, now has meaningfully different options to weigh, and meaningfully different trade-offs to understand.
What Staniszewski actually said about ElevenLabs' business
ElevenLabs says it is pacing at $600 million in annual recurring revenue, with 55 percent or more coming from classic enterprise customers including Klarna, Deutsche Telekom, Cisco, and Adobe. The remaining 45 percent is split across small businesses, developers, and creators. Staniszewski declined to discuss gross margins in detail, but he was explicit that he is willing to let them compress further if it means capturing more of the market. He also said that when ElevenLabs can pass savings on to customers, it does. On IPO timing, reported elsewhere as 2028, he would not confirm a date, saying only that the company is preparing the foundation to do it in the next few years. On whether callers should be told they are speaking to a bot, he said he thinks they should be told. That is a notable position for the CEO of a company whose voice models power first-line phone support for tens of millions of consumers who may not currently know it.
Google Gemini 3.8 TTS: what it ships and what it does not
Gemini 3.8 Flash TTS is positioned for creative direction, character design, gaming, audiobooks, and podcasts. It can generate new voices from natural language prompts and supports more than 100 languages and dialects. Gemini 3.8 Flash-Lite TTS targets audio dubbing, content creation, and voice agents. Analysts quoted by AI Business are measured about the release. Bradley Shimmin of Futurum Group described it as a refinement rather than a leap, useful for narrated long-form content and entertainment formats. Shimmin's case for it is integration: data professionals, he said, name integrating technologies as their biggest challenge with AI, and a voice model in the same family as the rest of Gemini makes that easier. Carter Huffman, CEO of AI voice company Modulate, agreed that one cohesive endpoint is a win, but only if it performs at the top of the field. He called Google's voice capabilities still nascent, and said Google would need a capability or application nobody else has to stand out.

The specialist-versus-platform trade-off buyers now face
ElevenLabs built its position as a specialist voice layer that other platforms plug into. It lets enterprise customers choose their own reasoning model from a menu, mixing frontier lab models with open-weight alternatives depending on the sensitivity and complexity of the task. Staniszewski noted that for purely informational customer service calls, open-source models work well because the knowledge base does most of the heavy lifting, but for financial services with authentication requirements, customers want the assurance of a frontier model. Google's argument with Gemini 3.8 TTS is the opposite: one cohesive endpoint that handles voice alongside everything else in the Gemini family, reducing integration overhead. For teams already building on Gemini, that saves a step. Teams that want to pick the voice and the reasoning model separately get that flexibility from a specialist. You can compare the options side by side in our AI audio tool reviews.
The disclosure question matters beyond ElevenLabs
Staniszewski's view that businesses should tell callers they are speaking to a bot is worth pausing on. ElevenLabs' models already power customer-facing voice interactions at significant scale, including for Klarna's 35 million US customers, and the lines between vendors are blurring: Decagon, a customer, trained its voice product on ElevenLabs and now competes with it. As more support calls are answered by a voice agent, the disclosure question stops being a philosophical nicety. Staniszewski expects it to change: in five years, he said, people will call in expecting an agent. For buyers evaluating voice AI tools, the practical implication is this: ask a vendor how its voice agents identify themselves on a call, and check that it fits what your customers expect.