Gemini 3.1 Flash Live
Gemini 3.1 Flash Live is a verified current AI model with official specifications, pricing or access details, capabilities,…
GPT-Realtime-2.1 is OpenAI's live speech reasoning model with 128K context, text/image/audio input, tool calls and improved interruption handling.
GPT-Realtime-2.1 is OpenAI's realtime reasoning model for voice agents that can listen, speak, see images, call tools and handle interruptions during live conversations.
GPT-Realtime-2.1 is OpenAI’s current reasoning-focused realtime model for applications where the user speaks naturally and expects the model to listen, respond, interrupt, resume and take actions without a slow transcription-chat-synthesis pipeline.
Traditional voice systems often chain speech-to-text, an LLM and text-to-speech. GPT-Realtime-2.1 instead handles audio directly and can produce audio directly, while still accepting text and image context. OpenAI says version 2.1 improves alphanumeric recognition, silence/noise handling and interruption behavior over GPT-Realtime-2. These details matter in real calls: account numbers, pauses, background noise and users talking over the assistant are common failure points that can make a voice agent feel unusable even when the underlying language intelligence is strong.
The model has a 128K context window and up to 32K output tokens. It accepts text, image and audio input, and returns text or audio. Reasoning is configurable and function calling is supported, allowing the voice agent to retrieve records, book appointments, update systems or hand off to other services during the conversation. Structured outputs are not supported, so tool schemas and server-side validation are especially important when actions affect real systems.
A customer-support agent can listen to a caller, check an account through a function, explain the result and continue naturally if the caller interrupts. A tutoring application can react to spoken questions and images of homework. A field-service agent can discuss a photo while the user speaks. SIP support also makes the model relevant to telephony, while WebRTC is appropriate for browser/mobile experiences and WebSocket connections fit server-mediated systems.
GPT-Realtime-2.1 is active through OpenAI’s realtime stack. Its knowledge cutoff is September 30, 2024, so current information should come from tools or application data rather than the model’s internal memory. The model is not fine-tunable and does not support structured outputs, but function calling makes it suitable for operational voice agents.
Text tokens cost $4/M input and $24/M output. Audio is much more expensive at $32/M input and $64/M output, with separate image-input pricing. Voice products should therefore optimize turn length, avoid unnecessary verbal repetition and design good endpointing/silence rules. A cheap text model may still be better for asynchronous processing behind the scenes, while Realtime should be reserved for the part of the experience that truly benefits from live speech.
Choose GPT-Realtime-2.1 when low-latency natural conversation, interruptions and tool use are central to the product. If you only need transcription, a dedicated transcription model will usually be cheaper. If you need a text chatbot with occasional generated audio, a text model plus TTS can be simpler. Realtime earns its cost when the conversation itself is the interface.
Realtime systems add operational complexity beyond model quality: audio devices fail, network jitter changes latency, people speak over each other and tools can return slowly or incorrectly. Applications should expose clear recovery behavior, confirm consequential actions and monitor both audio-token cost and end-to-end latency. Because the knowledge cutoff is older than current flagship text models, tool grounding is essential for fresh information.
Act as a live support agent, ask clarifying questions and call the account lookup tool when needed.Guide the user through this troubleshooting process while allowing interruptions.Look at this uploaded image and discuss it naturally in the voice conversation.Text: $4/M input, $0.40/M cached input and $24/M output. Audio: $32/M input and $64/M output. Image input is $5/M image tokens.
A strong OpenAI choice for serious voice agents because it combines low-latency conversation with reasoning and function calling rather than treating speech as a separate transcription step.