GPT-Audio-1.5
GPT-Audio-1.5 is a verified current AI model with official specifications, pricing or access details, capabilities, practical use cases, strengths and limitations.
What is this model and why does it matter?
GPT-Audio-1.5 is OpenAI's general-availability audio-in/audio-out model for Chat Completions, supporting spoken interactions, audio understanding and tool-enabled conversational workflows.
GPT-Audio-1.5: features, use cases and important details
GPT-Audio-1.5 is a current AI model verified from first-party OpenAI sources.
GPT-Audio-1.5 verified specifications
GPT-Audio-1.5 is OpenAI’s general-availability audio-in/audio-out model for Chat Completions, supporting spoken interactions, audio understanding and tool-enabled conversational workflows. Its verified context or usage limit is 128000 tokens, with 16384 tokens maximum output.
GPT-Audio-1.5 pricing and access
GPT-Audio-1.5 costs $2.50 per 1M text input tokens and $10 per 1M text output tokens. Audio costs $32 per 1M input audio tokens and $64 per 1M output audio tokens.
GPT-Audio-1.5 best uses
Voice assistants, Audio understanding, Spoken interfaces, Conversational applications, Function-enabled audio workflows.
GPT-Audio-1.5 limitations
No image or video input, Audio cost can dominate high-volume workloads, Voice workflows still need latency testing.
How to use this model
- Create an OpenAI API key.
- Call gpt-audio-1.5 through Chat Completions.
- Provide text and/or audio input.
- Request text or audio output.
- Add function tools where needed.
- Track audio-token usage separately from text.
Example prompts
Listen to this audio and answer the user's question.Respond to the user naturally with spoken audio.Use the available function after understanding the spoken request.
What it can do
- 128K context
- 16K output
- Audio input/output
- Text input/output
- Function calling
- Streaming
Practical use cases
- Voice bots
- Audio Q&A
- Accessibility
- Spoken customer support
- Interactive assistants
What does it cost?
GPT-Audio-1.5 costs $2.50 per 1M text input tokens and $10 per 1M text output tokens. Audio costs $32 per 1M input audio tokens and $64 per 1M output audio tokens.
What stands out
- GA audio model
- Audio in and out
- Function calling
- Large context
Things to consider
- Audio tokens are expensive
- No structured outputs
- No fine-tuning
Important restrictions and trade-offs
- No image or video input
- Audio cost can dominate high-volume workloads
- Voice workflows still need latency testing
Our editorial take
A strong current OpenAI audio model for developers who need flexible audio input and output without using the Realtime-specific model family.