Speech
Last updated:2026-08-12· 14 min read
🚀 Quick access
- ChatGPT Domestic:Open entry↗
- Mirror site:Open mirror↗
- Official ChatGPT:chatgpt.com ↗

Last updated: 2026-08-12
Overview
Speech is the “last mile” for many AI products: turn spoken input into searchable text, then read model replies aloud. OpenAI Platform offers Speech-to-Text (STT) and Text-to-Speech (TTS) REST endpoints, combinable with Realtime API and multimodal Responses. Model names (e.g. whisper-1, gpt-4o-transcribe, tts-1, tts-1-hd), supported formats, and pricing follow Speech docs and pricing—verify before launch.
Architecture: REST or Realtime?
| Mode | Best for | Latency | Complexity |
|---|---|---|---|
| REST STT + Responses + TTS | Upload recordings, meeting notes, offline batch | Seconds to tens of seconds | Low, easy to debug |
| Realtime full duplex | Voice assistant, interpreter-style interaction | Sub-second to seconds | High, WebSocket |
| STT only | Subtitles, compliance archive, search index | — | Lowest |
| TTS only | Announcements, accessibility read-aloud | — | Lowest |
Advice: ship REST three-stage loop first; evaluate Realtime when latency ROI justifies engineering cost.
STT: transcription endpoints
| Endpoint | Purpose |
|---|---|
/v1/audio/transcriptions | Transcribe in source language |
/v1/audio/translations | Non-English audio → English (per docs) |
curl https://api.openai.com/v1/audio/transcriptions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F file="@call.mp3" \
-F model="whisper-1" \
-F language="en" \
-F response_format="verbose_json"
from openai import OpenAI
client = OpenAI()
with open("call.mp3", "rb") as f:
tx = client.audio.transcriptions.create(
model="whisper-1", # verify current ID on Models page
file=f,
language="en",
response_format="verbose_json",
)
print(tx.text)
Preprocessing and splitting
| Practice | Why |
|---|---|
| Formats listed in docs (mp3, wav, m4a) | Fewer codec failures |
| Split long audio on silence or 5–10 min chunks | Avoid size/timeout limits |
| Mono, light denoise | Better recognition |
Pass language="en" (or target) | Reduces mis-detection |
Normalize punctuation and filter sensitive terms before downstream use; summarization and classification belong in Responses, not crammed into STT params.
Output formats
text: simplest plain textverbose_json: segments with timestamps for player seeksrt/vtt: subtitle files—verify player compatibility
TTS: synthesis endpoint
POST /v1/audio/speech converts text to mp3, opus, or pcm.
from pathlib import Path
from openai import OpenAI
client = OpenAI()
speech = client.audio.speech.create(
model="tts-1",
voice="nova",
input="Your ticket has been received. We will reply within two business days.",
response_format="mp3",
)
Path("reply.mp3").write_bytes(speech.content)
| Parameter | Dev notes |
|---|---|
voice | List per docs; pick 1–2 for product consistency |
speed | If supported, IVR may use slight speed-up |
| stream | Long read-aloud: generate while playing, lower time-to-first-byte |
Caching: fixed phrases (verification codes, standard replies) by (voice, text) hash can cut cost sharply.
Typical pipeline
User recording → object storage → STT → text cleanup → Responses (summary / intent / reply)
↓
optional TTS → CDN → client playback
Combine with Vision: video audio STT + key-frame vision. Index transcripts via Embeddings guide after PII redaction.
Billing, rate limits, and compliance
- STT often bills by minute; TTS by characters—see openai.com/api/pricing
- Exponential backoff on 429 / 5xx; batch via async workers
- Recording transcription needs user consent and privacy disclosure
- Medical/legal: human review; don’t TTS impersonate others without authorization
Frequently asked questions
Whisper misses domain jargon?
Set language, post-process with domain glossary in Responses; extreme cases add custom dictionary + spot checks.
File exceeds size limit?
Split on silence or fixed duration; check Batch API in docs.
Can TTS clone a real person?
Follow platform policy; commercial voice cloning needs legal authorization.
Mix REST and Realtime?
Yes: Realtime for live dialog, REST STT for offline archive, shared Responses backend.
Store transcripts directly?
Minimize retention, encrypt, set TTL; redact before DB or index.
Official resources
Next reading
Action path
Today: Transcribe a 1-minute mp3; save verbose_json. Tomorrow: TTS the same reply text; compare tts-1 vs hd tier. This week: Build upload → STT → Responses summary → optional TTS; log STT cost per minute.
Related
OpenAI Dev Overview
2026 OpenAI developer map: how ChatGPT web, Platform console, and APIs divide work—and the reading order from first call to production.
OpenAI Platform Overview
platform.openai.com console, doc navigation, Playground, usage billing, and org management—how developers find API information efficiently.
OpenAI API Quickstart
From Platform account and API key to your first OpenAI call: Responses/Completions examples, billing, rate limits, and a security checklist (2026 hands-on).
ChatGPT API Developer Guide
Production OpenAI API integration: architecture, auth, streaming, tool use, rate-limit retries, and a launch checklist.