Sponsored by Byond Boundrys Consulting - Empowering Ideas, Delivering Results
OpenAI Intermediate

GPT-Audio-1.5

GPT-Audio-1.5 is a verified current AI model with official specifications, pricing or access details, capabilities, practical use cases, strengths and limitations.

Audio Language ModelTextAudio Paid
In plain English

What is this model and why does it matter?

GPT-Audio-1.5 is OpenAI's general-availability audio-in/audio-out model for Chat Completions, supporting spoken interactions, audio understanding and tool-enabled conversational workflows.

Voice assistantsAudio understandingSpoken interfacesConversational applicationsFunction-enabled audio workflows
Model overview

GPT-Audio-1.5: features, use cases and important details

GPT-Audio-1.5 is a current AI model verified from first-party OpenAI sources.

GPT-Audio-1.5 verified specifications

GPT-Audio-1.5 is OpenAI’s general-availability audio-in/audio-out model for Chat Completions, supporting spoken interactions, audio understanding and tool-enabled conversational workflows. Its verified context or usage limit is 128000 tokens, with 16384 tokens maximum output.

GPT-Audio-1.5 pricing and access

GPT-Audio-1.5 costs $2.50 per 1M text input tokens and $10 per 1M text output tokens. Audio costs $32 per 1M input audio tokens and $64 per 1M output audio tokens.

GPT-Audio-1.5 best uses

Voice assistants, Audio understanding, Spoken interfaces, Conversational applications, Function-enabled audio workflows.

GPT-Audio-1.5 limitations

No image or video input, Audio cost can dominate high-volume workloads, Voice workflows still need latency testing.

Get started

How to use this model

  1. Create an OpenAI API key.
  2. Call gpt-audio-1.5 through Chat Completions.
  3. Provide text and/or audio input.
  4. Request text or audio output.
  5. Add function tools where needed.
  6. Track audio-token usage separately from text.
Copy and try

Example prompts

  • Listen to this audio and answer the user's question.
  • Respond to the user naturally with spoken audio.
  • Use the available function after understanding the spoken request.
Capabilities

What it can do

  • 128K context
  • 16K output
  • Audio input/output
  • Text input/output
  • Function calling
  • Streaming
Best for

Practical use cases

  • Voice bots
  • Audio Q&A
  • Accessibility
  • Spoken customer support
  • Interactive assistants
Pricing

What does it cost?

GPT-Audio-1.5 costs $2.50 per 1M text input tokens and $10 per 1M text output tokens. Audio costs $32 per 1M input audio tokens and $64 per 1M output audio tokens.

Input$2.50/M text; $32/M audio
Output$10/M text; $64/M audio
Simple summaryText and audio use separate token rates, with audio substantially more expensive than text.

What stands out

  • GA audio model
  • Audio in and out
  • Function calling
  • Large context

Things to consider

  • Audio tokens are expensive
  • No structured outputs
  • No fine-tuning
Limitations

Important restrictions and trade-offs

  • No image or video input
  • Audio cost can dominate high-volume workloads
  • Voice workflows still need latency testing
SimplifyAITools verdict

Our editorial take

A strong current OpenAI audio model for developers who need flexible audio input and output without using the Realtime-specific model family.

References

Primary sources

  1. Open source 1 ↗
  2. Open source 2 ↗