top of page

Gemini 3.8 Live and Extended Thinking: The Technology Powering the Next Wave of AI Agents

23 hours ago
9 min read
Google’s latest Gemini Audio models signal an important shift in the development of conversational artificial intelligence. With the introduction of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, voice AI is moving beyond the traditional model of listening, transcribing, generating an answer, and speaking it back. The emerging architecture is designed to reason, interact with visual information, invoke tools, and perform multi-step tasks while maintaining an uninterrupted conversation.

This distinction matters because natural voice interaction is fundamentally different from text-based prompting. Humans expect dialogue to remain continuous. They interrupt, change topics, provide incomplete information, point at objects, correct themselves, switch languages, and expect an assistant to continue working while they talk. Google’s new models are designed around precisely this type of interaction.

The broader significance extends beyond consumer assistants. Real-time multimodal voice systems could become an important interface for enterprise software, customer service, education, robotics, financial services, accessibility tools, automotive systems, and agentic applications.

From Speech Recognition to Conversational Agents

Earlier generations of voice assistants generally relied on a pipeline consisting of automatic speech recognition, language processing, task execution, and text-to-speech generation. Each component performed a distinct function. While this architecture remains useful, it can introduce latency and make conversations feel mechanical.

Native speech-to-speech systems approach the problem differently. Instead of treating voice primarily as a transcription layer, they can process audio as part of a richer conversational context and generate spoken responses directly.

Gemini 3.8 Live extends this concept by combining speech interaction with reasoning, visual grounding, and tool execution. A user can therefore speak to an agent while simultaneously providing visual information, asking it to perform an action, or changing the direction of the conversation.

The practical difference is substantial. An assistant that merely answers questions is reactive. An agent capable of maintaining dialogue while executing tasks becomes collaborative.

That distinction is central to the evolution of voice AI.

Gemini 3.8 Live vs. Gemini 3.8 Live Extended Thinking

Google has positioned the two models for different operational requirements.

Capability	Gemini 3.8 Live	Gemini 3.8 Live Extended Thinking
Primary focus	Scalable, fluid real-time dialogue	Complex reasoning and task completion
Voice interaction	Real-time speech-to-speech	Real-time speech-to-speech with deeper reasoning
Visual understanding	Near-real-time visual context	Visual context combined with deeper reasoning
Tool execution	Background API and tool calls	Background multi-step execution
Reasoning	Conversational reasoning	Configurable, extended reasoning
Ideal applications	Customer interaction, assistants, live support	Complex workflows, research, planning, agentic tasks
Conversational continuity	Designed for uninterrupted dialogue	Designed to reason while continuing the conversation

Gemini 3.8 Live is therefore particularly relevant where responsiveness, scale, and operating cost are priorities. Extended Thinking targets situations in which the agent must solve more complicated problems or coordinate several actions.

This distinction resembles the broader evolution of AI models toward specialized inference strategies. Not every request requires maximum reasoning depth. Simple interactions benefit from speed, while complicated workflows can justify additional computation.

Reasoning While Speaking Changes the User Experience

One of the most consequential capabilities in Gemini 3.8 Live Extended Thinking is its ability to reason while maintaining spoken interaction.

Traditional AI systems often become silent during a lengthy operation. The user waits for an answer without knowing whether the system is processing a request, encountering an error, or simply stalled.

Extended Thinking addresses this interaction problem through conversational progress signals. An agent can acknowledge a request early, continue processing in the background, and communicate progress as a complex task develops.

This is more than a cosmetic improvement. Human conversations contain continuous feedback. Short acknowledgments establish that the other participant has heard and understood the request. Applying this principle to AI can reduce perceived latency and make longer-running agentic workflows more understandable.

The result is a different interaction model: the user does not necessarily wait for the machine to finish thinking before conversation continues.

Visual Context Makes Voice AI Multimodal

Voice becomes substantially more useful when it can be connected to visual information.

Gemini 3.8 Live can process visual inputs in near real time, allowing an interaction to incorporate what a user is seeing alongside what the user is saying. This creates possibilities that conventional voice assistants cannot easily support.

Consider a technical support scenario. Instead of describing an unfamiliar device, a user could show the problem through a camera while verbally explaining what happened. The AI could combine the visual evidence and spoken description when determining what assistance to provide.

The same principle applies to education, workplace training, accessibility, retail, gaming, and field operations.

A worker could receive spoken guidance while showing equipment to an AI assistant. A student could point a camera toward a diagram and ask for an explanation. A customer could display a product while asking how to configure it.

The important development is not simply that AI can see. It is that visual understanding can become part of a continuous conversational state.

Background Tool Execution Moves AI Toward Agents

Another major development is asynchronous function calling.

In a conventional conversational system, an external operation can interrupt the interaction. The assistant may need to wait for a database query, API response, booking operation, or other tool call before continuing.

Asynchronous execution allows the model to initiate those operations while continuing to communicate.

This architecture is particularly important for agentic AI. Real-world tasks rarely consist of a single question and a single answer. They frequently require multiple steps, external systems, verification, and follow-up actions.

A voice agent could, for example, gather information from a user, initiate an external operation, ask another relevant question while waiting, and incorporate the returned information when it becomes available.

The underlying concept is similar to parallelism in computing. Instead of forcing every operation into a strictly sequential conversational loop, independent processes can proceed concurrently.

For enterprises, this can translate into more efficient customer interactions and potentially shorter perceived task times.

Multilingual Conversation Becomes More Practical

Google reports that Gemini 3.8 Live can automatically detect and transition among 97 supported languages during conversations.

Multilingual interaction is especially important for voice because language switching frequently occurs naturally. Users may move between languages within a conversation, use regional expressions, pronounce technical terminology differently, or communicate with accents that vary substantially from standardized training examples.

Automatic code-switching can reduce the need for users to manually configure language settings. For international businesses, this could make a single conversational interface practical across diverse customer populations.

The dedicated Gemini 3.5 Transcribe model addresses the complementary speech-to-text problem. Google reports support for more than 85 languages, with reported Word Error Rates of 4.0% for streaming transcription and 2.6% for non-streaming transcription.

Its capabilities include automatic code-switching, custom vocabulary support, and a smart transcription mode designed to produce cleaner, structured text.

These features are valuable because transcription is not merely a technical conversion. In professional environments, specialized terminology, names, product identifiers, claim numbers, and codes can determine whether a voice system is genuinely useful.

Why Alphanumeric Precision Matters

Voice interfaces have historically struggled with information such as confirmation codes, account identifiers, technical specifications, serial numbers, and claim references.

A conversational model can be highly intelligent yet still fail if it misunderstands a single character in a code.

Google highlights alphanumeric precision as a capability of its new Live models. This is particularly relevant to banking, insurance, logistics, telecommunications, healthcare administration, technical support, and enterprise operations.

For these applications, accuracy cannot be measured only by how natural the conversation sounds. The system must also reliably preserve operational information.

This illustrates a broader principle for enterprise AI: conversational quality and transactional accuracy are separate dimensions, and production systems require both.

The Developer Ecosystem Is Becoming More Important

The emergence of real-time voice agents also changes the software stack required to deploy them.

Developers need more than an AI model. Production voice applications require real-time media transport, session management, audio processing, authentication, tool integration, monitoring, and scalable infrastructure.

Google’s Live API ecosystem includes integrations with platforms such as Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents. These platforms can abstract portions of the underlying media infrastructure, allowing developers to concentrate more heavily on application logic and user experience.

The importance of this ecosystem should not be underestimated. Historically, infrastructure complexity has often slowed the adoption of emerging computing interfaces. Making real-time multimodal AI easier to integrate could accelerate experimentation and commercial deployment.

Business Impact: Voice Becomes an Interface for Work

The most significant commercial opportunity may not be conversational entertainment. It may be the transformation of voice into an operational interface.

Employees could interact with enterprise systems without navigating multiple screens. Customer-service agents could receive contextual assistance during live conversations. Sales teams could use voice to retrieve information and prepare materials. Field technicians could obtain instructions while working with equipment. Training platforms could simulate realistic interactions with employees.

Google’s demonstrations point toward this broader direction, including onboarding, customer interactions, document generation, business planning, technical troubleshooting, and asynchronous reservations.

The common theme is that AI is becoming capable of performing work rather than simply describing how work should be performed.

That creates an important economic distinction. A chatbot primarily produces information. An AI agent can potentially produce outcomes.

The Challenges Behind Real-Time Voice AI

The technology also introduces significant challenges.

Latency remains critical. A system that is technically intelligent but responds too slowly can feel unnatural. Audio quality, network reliability, interruption handling, and turn-taking all influence perceived intelligence.

Reliability is another concern. When voice agents are connected to external systems, errors can become operational rather than merely conversational. An incorrect interpretation could result in a failed transaction, inaccurate customer information, or an inappropriate action.

Privacy and security also become more important as AI systems process continuous audio and visual information. Organizations deploying voice agents must carefully manage authentication, data retention, permissions, monitoring, and access to connected tools.

There is also a fundamental challenge in evaluating conversational agents. Traditional language benchmarks do not fully capture whether an AI system can maintain a natural conversation while completing a real-world task. Voice-agent evaluation increasingly needs to measure reasoning, latency, task completion, interruption handling, speech quality, and user experience together.

The Economics of Real-Time AI

Cost is another critical factor in determining adoption.

Google reports pricing for Gemini 3.8 Live at $0.005 per minute for audio input and $0.018 per minute for audio output. For developers, this makes inference economics an important design consideration, particularly for applications handling millions of minutes of interaction.

The distinction between Gemini 3.8 Live and Extended Thinking also highlights a broader industry trend toward intelligent allocation of computational resources.

A production system does not necessarily need maximum reasoning for every conversation. Organizations can potentially reserve deeper reasoning for complex tasks while using faster inference for routine interactions.

This creates the possibility of dynamically matching computational intensity to task complexity.

What Comes Next for Voice Agents

The next phase of voice AI is likely to involve increasingly autonomous systems that combine continuous perception, reasoning, memory, external tools, and multimodal interaction.

The progression can be viewed as four broad stages:

Speech recognition, converting spoken language into text.
Conversational AI, generating intelligent responses to spoken requests.
Agentic voice AI, performing actions through tools and external systems.
Multimodal autonomous assistants, continuously interpreting voice, vision, context, and ongoing tasks.

Gemini 3.8 Live and Extended Thinking sit firmly within the third stage while providing infrastructure for movement toward the fourth.

The long-term competitive question will therefore not be which AI can simply speak most naturally. It will be which systems can combine intelligence, speed, reliability, contextual awareness, and safe action execution into a coherent interface.

Conclusion

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking represent a broader transition in artificial intelligence, from systems that respond to prompts toward systems that participate in ongoing activities.

Their significance lies in the combination of real-time speech, visual understanding, asynchronous tools, multilingual interaction, and deeper reasoning. Gemini 3.5 Transcribe complements this ecosystem by addressing high-precision speech recognition for applications where accurate transcription is essential.

The most important development is the changing role of voice itself. Voice is increasingly becoming more than an input mechanism. It is evolving into an interface through which people can access information, coordinate software, complete workflows, and collaborate with intelligent agents.

For businesses, developers, and researchers, this shift deserves close attention. As latency falls, multimodal reasoning improves, and agent frameworks mature, real-time voice AI could become one of the most consequential interfaces in the next generation of computing.

For technology analysts such as Dr. Shahid Masood and the expert team at 1950.ai, the development illustrates a broader transformation already underway across artificial intelligence: the convergence of reasoning, multimodality, automation, and human-computer interaction. The future of AI may not simply be about building models that know more. It may increasingly be about building systems that can understand what people are doing, participate naturally, and safely help accomplish what comes next.

Key Takeaways
Gemini 3.8 Live focuses on scalable, low-latency conversational AI with visual and tool-based capabilities.
Gemini 3.8 Live Extended Thinking is designed for more complex reasoning and multi-step agentic workflows.
Asynchronous function calling allows tools and APIs to operate while dialogue continues.
Near-real-time visual processing expands voice AI into genuinely multimodal interaction.
Support for 97 languages strengthens the potential for global conversational applications.
Gemini 3.5 Transcribe targets high-precision speech-to-text across more than 85 languages.
Real-world adoption will depend on reliability, latency, security, privacy, evaluation quality, and inference economics.
The larger industry trend is the transformation of voice from a simple command interface into an operational layer for AI agents.
Further Reading / External References

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/

Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe

https://blog.google/innovation-and-ai/technology/developers-tools/build-real-time-voice-applications-gemini-audio/

Google Launches Gemini 3.8 Live and Extended Thinking Voice Models

https://www.unite.ai/google-launches-gemini-3-8-live-and-extended-thinking-voice-models/

Google’s latest Gemini Audio models signal an important shift in the development of conversational artificial intelligence. With the introduction of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, voice AI is moving beyond the traditional model of listening, transcribing, generating an answer, and speaking it back. The emerging architecture is designed to reason, interact with visual information, invoke tools, and perform multi-step tasks while maintaining an uninterrupted conversation.


This distinction matters because natural voice interaction is fundamentally different from text-based prompting. Humans expect dialogue to remain continuous. They interrupt, change topics, provide incomplete information, point at objects, correct themselves, switch languages, and expect an assistant to continue working while they talk. Google’s new models are designed around precisely this type of interaction.

The broader significance extends beyond consumer assistants. Real-time multimodal voice systems could become an important interface for enterprise software, customer service, education, robotics, financial services, accessibility tools, automotive systems, and agentic applications.


From Speech Recognition to Conversational Agents

Earlier generations of voice assistants generally relied on a pipeline consisting of automatic speech recognition, language processing, task execution, and text-to-speech generation. Each component performed a distinct function. While this architecture remains useful, it can introduce latency and make conversations feel mechanical.

Native speech-to-speech systems approach the problem differently. Instead of treating voice primarily as a transcription layer, they can process audio as part of a richer conversational context and generate spoken responses directly.


Gemini 3.8 Live extends this concept by combining speech interaction with reasoning, visual grounding, and tool execution. A user can therefore speak to an agent while simultaneously providing visual information, asking it to perform an action, or changing the direction of the conversation.

The practical difference is substantial. An assistant that merely answers questions is reactive. An agent capable of maintaining dialogue while executing tasks becomes collaborative.

That distinction is central to the evolution of voice AI.


Gemini 3.8 Live vs. Gemini 3.8 Live Extended Thinking

Google has positioned the two models for different operational requirements.

Capability

Gemini 3.8 Live

Gemini 3.8 Live Extended Thinking

Primary focus

Scalable, fluid real-time dialogue

Complex reasoning and task completion

Voice interaction

Real-time speech-to-speech

Real-time speech-to-speech with deeper reasoning

Visual understanding

Near-real-time visual context

Visual context combined with deeper reasoning

Tool execution

Background API and tool calls

Background multi-step execution

Reasoning

Conversational reasoning

Configurable, extended reasoning

Ideal applications

Customer interaction, assistants, live support

Complex workflows, research, planning, agentic tasks

Conversational continuity

Designed for uninterrupted dialogue

Designed to reason while continuing the conversation

Gemini 3.8 Live is therefore particularly relevant where responsiveness, scale, and operating cost are priorities. Extended Thinking targets situations in which the agent must solve more complicated problems or coordinate several actions.

This distinction resembles the broader evolution of AI models toward specialized inference strategies. Not every request requires maximum reasoning depth. Simple interactions benefit from speed, while complicated workflows can justify additional computation.


Reasoning While Speaking Changes the User Experience

One of the most consequential capabilities in Gemini 3.8 Live Extended Thinking is its ability to reason while maintaining spoken interaction.

Traditional AI systems often become silent during a lengthy operation. The user waits for an answer without knowing whether the system is processing a request, encountering an error, or simply stalled.

Extended Thinking addresses this interaction problem through conversational progress signals. An agent can acknowledge a request early, continue processing in the background, and communicate progress as a complex task develops.


This is more than a cosmetic improvement. Human conversations contain continuous feedback. Short acknowledgments establish that the other participant has heard and understood the request. Applying this principle to AI can reduce perceived latency and make longer-running agentic workflows more understandable.

The result is a different interaction model: the user does not necessarily wait for the machine to finish thinking before conversation continues.


Visual Context Makes Voice AI Multimodal

Voice becomes substantially more useful when it can be connected to visual information.

Gemini 3.8 Live can process visual inputs in near real time, allowing an interaction to incorporate what a user is seeing alongside what the user is saying. This creates possibilities that conventional voice assistants cannot easily support.

Consider a technical support scenario. Instead of describing an unfamiliar device, a user could show the problem through a camera while verbally explaining what happened. The AI could combine the visual evidence and spoken description when determining what assistance to provide.


The same principle applies to education, workplace training, accessibility, retail, gaming, and field operations.

A worker could receive spoken guidance while showing equipment to an AI assistant. A student could point a camera toward a diagram and ask for an explanation. A customer could display a product while asking how to configure it.

The important development is not simply that AI can see. It is that visual understanding can become part of a continuous conversational state.


Background Tool Execution Moves AI Toward Agents

Another major development is asynchronous function calling.

In a conventional conversational system, an external operation can interrupt the interaction. The assistant may need to wait for a database query, API response, booking operation, or other tool call before continuing.

Asynchronous execution allows the model to initiate those operations while continuing to communicate.

This architecture is particularly important for agentic AI. Real-world tasks rarely consist of a single question and a single answer. They frequently require multiple steps, external systems, verification, and follow-up actions.

A voice agent could, for example, gather information from a user, initiate an external operation, ask another relevant question while waiting, and incorporate the returned information when it becomes available.


The underlying concept is similar to parallelism in computing. Instead of forcing every operation into a strictly sequential conversational loop, independent processes can proceed concurrently.

For enterprises, this can translate into more efficient customer interactions and potentially shorter perceived task times.


Multilingual Conversation Becomes More Practical

Google reports that Gemini 3.8 Live can automatically detect and transition among 97 supported languages during conversations.

Multilingual interaction is especially important for voice because language switching frequently occurs naturally. Users may move between languages within a conversation, use regional expressions, pronounce technical terminology differently, or communicate with accents that vary substantially from standardized training examples.

Automatic code-switching can reduce the need for users to manually configure language settings. For international businesses, this could make a single conversational interface practical across diverse customer populations.


The dedicated Gemini 3.5 Transcribe model addresses the complementary speech-to-text problem. Google reports support for more than 85 languages, with reported Word Error Rates of 4.0% for streaming transcription and 2.6% for non-streaming transcription.

Its capabilities include automatic code-switching, custom vocabulary support, and a smart transcription mode designed to produce cleaner, structured text.

These features are valuable because transcription is not merely a technical conversion. In professional environments, specialized terminology, names, product identifiers, claim numbers, and codes can determine whether a voice system is genuinely useful.


Why Alphanumeric Precision Matters

Voice interfaces have historically struggled with information such as confirmation codes, account identifiers, technical specifications, serial numbers, and claim references.

A conversational model can be highly intelligent yet still fail if it misunderstands a single character in a code.

Google highlights alphanumeric precision as a capability of its new Live models. This is particularly relevant to banking, insurance, logistics, telecommunications, healthcare administration, technical support, and enterprise operations.


For these applications, accuracy cannot be measured only by how natural the conversation sounds. The system must also reliably preserve operational information.

This illustrates a broader principle for enterprise AI: conversational quality and transactional accuracy are separate dimensions, and production systems require both.


The Developer Ecosystem Is Becoming More Important

The emergence of real-time voice agents also changes the software stack required to deploy them.

Developers need more than an AI model. Production voice applications require real-time media transport, session management, audio processing, authentication, tool integration, monitoring, and scalable infrastructure.


Google’s Live API ecosystem includes integrations with platforms such as Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents. These platforms can abstract portions of the underlying media infrastructure, allowing developers to concentrate more heavily on application logic and user experience.

The importance of this ecosystem should not be underestimated. Historically, infrastructure complexity has often slowed the adoption of emerging computing interfaces. Making real-time multimodal AI easier to integrate could accelerate experimentation and commercial deployment.


Business Impact: Voice Becomes an Interface for Work

The most significant commercial opportunity may not be conversational entertainment. It may be the transformation of voice into an operational interface.

Employees could interact with enterprise systems without navigating multiple screens. Customer-service agents could receive contextual assistance during live conversations. Sales teams could use voice to retrieve information and prepare materials. Field technicians could obtain instructions while working with equipment. Training platforms could simulate realistic interactions with employees.


Google’s demonstrations point toward this broader direction, including onboarding, customer interactions, document generation, business planning, technical troubleshooting, and asynchronous reservations.

The common theme is that AI is becoming capable of performing work rather than simply describing how work should be performed.

That creates an important economic distinction. A chatbot primarily produces information. An AI agent can potentially produce outcomes.


The Challenges Behind Real-Time Voice AI

The technology also introduces significant challenges.

Latency remains critical. A system that is technically intelligent but responds too slowly can feel unnatural. Audio quality, network reliability, interruption handling, and turn-taking all influence perceived intelligence.

Reliability is another concern. When voice agents are connected to external systems, errors can become operational rather than merely conversational. An incorrect interpretation could result in a failed transaction, inaccurate customer information, or an inappropriate action.


Privacy and security also become more important as AI systems process continuous audio and visual information. Organizations deploying voice agents must carefully manage authentication, data retention, permissions, monitoring, and access to connected tools.

There is also a fundamental challenge in evaluating conversational agents. Traditional language benchmarks do not fully capture whether an AI system can maintain a natural conversation while completing a real-world task. Voice-agent evaluation increasingly needs to measure reasoning, latency, task completion, interruption handling, speech quality, and user experience together.


The Economics of Real-Time AI

Cost is another critical factor in determining adoption.

Google reports pricing for Gemini 3.8 Live at $0.005 per minute for audio input and $0.018 per minute for audio output. For developers, this makes inference economics an important design consideration, particularly for applications handling millions of minutes of interaction.


The distinction between Gemini 3.8 Live and Extended Thinking also highlights a broader industry trend toward intelligent allocation of computational resources.

A production system does not necessarily need maximum reasoning for every conversation. Organizations can potentially reserve deeper reasoning for complex tasks while using faster inference for routine interactions.

This creates the possibility of dynamically matching computational intensity to task complexity.


What Comes Next for Voice Agents

The next phase of voice AI is likely to involve increasingly autonomous systems that combine continuous perception, reasoning, memory, external tools, and multimodal interaction.

The progression can be viewed as four broad stages:

  1. Speech recognition, converting spoken language into text.

  2. Conversational AI, generating intelligent responses to spoken requests.

  3. Agentic voice AI, performing actions through tools and external systems.

  4. Multimodal autonomous assistants, continuously interpreting voice, vision, context, and ongoing tasks.

Gemini 3.8 Live and Extended Thinking sit firmly within the third stage while providing infrastructure for movement toward the fourth.

The long-term competitive question will therefore not be which AI can simply speak most naturally. It will be which systems can combine intelligence, speed, reliability, contextual awareness, and safe action execution into a coherent interface.


Conclusion

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking represent a broader transition in artificial intelligence, from systems that respond to prompts toward systems that participate in ongoing activities.

Their significance lies in the combination of real-time speech, visual understanding, asynchronous tools, multilingual interaction, and deeper reasoning. Gemini 3.5 Transcribe complements this ecosystem by addressing high-precision speech recognition for applications where accurate transcription is essential.


The most important development is the changing role of voice itself. Voice is increasingly becoming more than an input mechanism. It is evolving into an interface through which people can access information, coordinate software, complete workflows, and collaborate with intelligent agents.

For businesses, developers, and researchers, this shift deserves close attention. As latency falls, multimodal reasoning improves, and agent frameworks mature, real-time voice AI could become one of the most consequential interfaces in the next generation of computing.


For technology analysts such as Dr. Shahid Masood and the expert team at 1950.ai, the development illustrates a broader transformation already underway across artificial intelligence: the convergence of reasoning, multimodality, automation, and human-computer interaction. The future of AI may not simply be about building models that know more. It may increasingly be about building systems that can understand what people are doing, participate naturally, and safely help accomplish what comes next.


Key Takeaways

  • Gemini 3.8 Live focuses on scalable, low-latency conversational AI with visual and tool-based capabilities.

  • Gemini 3.8 Live Extended Thinking is designed for more complex reasoning and multi-step agentic workflows.

  • Asynchronous function calling allows tools and APIs to operate while dialogue continues.

  • Near-real-time visual processing expands voice AI into genuinely multimodal interaction.

  • Support for 97 languages strengthens the potential for global conversational applications.

  • Gemini 3.5 Transcribe targets high-precision speech-to-text across more than 85 languages.

  • Real-world adoption will depend on reliability, latency, security, privacy, evaluation quality, and inference economics.

  • The larger industry trend is the transformation of voice from a simple command interface into an operational layer for AI agents.


Further Reading / External References

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe

Google Launches Gemini 3.8 Live and Extended Thinking Voice Models

Comments


bottom of page