GPT-Realtime-2.1 Mini
GPT-Realtime-2.1 Mini is a verified current AI model with official specifications, pricing or access information, capabilities, practical use…
Gemini 3.1 Flash Live is a verified current AI model with official specifications, pricing or access details, capabilities, practical use cases, strengths and limitations.
Gemini 3.1 Flash Live is Google's low-latency audio-to-audio preview model for realtime dialogue, voice-first agents and multimodal conversational experiences.
Gemini 3.1 Flash Live is a current AI model verified from first-party Google sources.
Gemini 3.1 Flash Live is Google’s low-latency audio-to-audio preview model for realtime dialogue, voice-first agents and multimodal conversational experiences. Its verified context or usage limit is 131072 input tokens, with 65536 output tokens maximum output.
Gemini 3.1 Flash Live has a free tier. Paid rates are $0.75/M text input, $3/M audio input (about $0.005/min), $1/M image/video input, $4.50/M text output and $12/M audio output (about $0.018/min).
Voice agents, Realtime assistants, Multimodal conversations, Customer support, Interactive applications.
Async function calling is not supported, Live sessions can accumulate context cost, Preview behavior can change.
In addition, its main capabilities include 131K input, 65K output, Realtime audio in/out, Image/video awareness, Function calling and Search grounding. For example, common use cases include Voice bots, Realtime tutoring, Contact centers, Multimodal assistants and Interactive apps.
In practice, this model may suit Voice agents, Realtime assistants, Multimodal conversations, Customer support and Interactive applications. Also, notable strengths include Low-latency audio-to-audio, Multimodal, Function calling and Free tier. However, review trade-offs such as Async function calling is not supported, Live sessions can accumulate context cost and Preview behavior can change before adopting it.
Meanwhile, Gemini 3.1 Flash Live has a free tier. Paid rates are $0.75/M text input, $3/M audio input (about $0.005/min), $1/M image/video input, $4.50/M text output and $12/M audio output (about $0.018/min). Free tier available; paid billing is modality-based and Live API context is reprocessed across turns.
Use the official model website, official documentation, pricing or release source and additional primary source to confirm current availability, limits and pricing. Product details can change after publication, so rely on primary documentation for final decisions.
Next, continue your research in the AI models directory, Google models and Realtime Voice Model models. Compare providers, pricing, modalities and practical limitations side by side to choose the right model for your workflow.
Act as a realtime voice assistant for this support call.Discuss this live video feed with the user.Listen to the speaker and use the available function when confirmation is given.Gemini 3.1 Flash Live has a free tier. Paid rates are $0.75/M text input, $3/M audio input (about $0.005/min), $1/M image/video input, $4.50/M text output and $12/M audio output (about $0.018/min).
A strong current voice-model target for users building Gemini realtime conversational and multimodal agents.