GPT Audio
Provider:OpenAI
The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...
Key specs
- Context:128K
- Max output:16.4K
- Tokenizer:GPT
- Released:2026-01-19
Token pricing
- Input price:$2.50 / 1M tokens
- Output price:$10.00 / 1M tokens
- Blended price:$4.38 / 1M tokens
Modalities
Text, Audio
Use Cases
- Deep Reasoning & Analysis:Math proofs, logical reasoning, multi-step planning — ideal for research and financial analysis.(≈ $0.080/call)
- Voice Interaction:Text-to-speech (TTS) with multiple voices and speed control — podcasts, audiobooks, IVR systems.
- Chat & Support:FAQ auto-response, multi-turn dialogue, emotion detection — e-commerce, finance, gov.(≈ $0.020/call)
- Content Moderation:Auto-detect violations, hate speech, NSFW content; custom rule sets and score thresholds.
When to pick this model
Signals derived from public pricing and spec fields — not a benchmark. Validate with official docs and your own eval.
- Audio support: fit for speech-to-text / voice agent / transcript analysis.
FAQ
- What is the token pricing of GPT Audio? Input price $2.50 / 1M tokens, output price $10.00 / 1M tokens, blended about $4.38 / 1M tokens. Refer to OpenAI’s official page for the exact rate.
- What context window does GPT Audio support? Context window is 128K, max output about 16.4K, suitable for long documents, multi-turn chat and complex reasoning.
- What scenarios is GPT Audio best for? Based on its capability and pricing, GPT Audio fits best: Deep Reasoning & Analysis, Voice Interaction, Chat & Support. See the Use Cases section for details and cost estimates.
- What input/output modalities does GPT Audio support? Supports Text, Audio modality.
- Is GPT Audio a free model? GPT Audio is billed per token, not a free model.
- Which provider offers GPT Audio? GPT Audio is offered by OpenAI.
More models from OpenAI
Full info:GPT Audio