Skip to content

Speech

Last updated:2026-08-12· 14 min read

🚀 Quick access

  • ChatGPT Domestic:Open entry↗
  • Mirror site:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

Speech

Last updated: 2026-08-12

Overview

Speech is the “last mile” for many AI products: turn spoken input into searchable text, then read model replies aloud. OpenAI Platform offers Speech-to-Text (STT) and Text-to-Speech (TTS) REST endpoints, combinable with Realtime API and multimodal Responses. Model names (e.g. whisper-1, gpt-4o-transcribe, tts-1, tts-1-hd), supported formats, and pricing follow Speech docs and pricing—verify before launch.

Architecture: REST or Realtime?

ModeBest forLatencyComplexity
REST STT + Responses + TTSUpload recordings, meeting notes, offline batchSeconds to tens of secondsLow, easy to debug
Realtime full duplexVoice assistant, interpreter-style interactionSub-second to secondsHigh, WebSocket
STT onlySubtitles, compliance archive, search index—Lowest
TTS onlyAnnouncements, accessibility read-aloud—Lowest

Advice: ship REST three-stage loop first; evaluate Realtime when latency ROI justifies engineering cost.

STT: transcription endpoints

EndpointPurpose
/v1/audio/transcriptionsTranscribe in source language
/v1/audio/translationsNon-English audio → English (per docs)
curl https://api.openai.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F file="@call.mp3" \
  -F model="whisper-1" \
  -F language="en" \
  -F response_format="verbose_json"
from openai import OpenAI

client = OpenAI()
with open("call.mp3", "rb") as f:
    tx = client.audio.transcriptions.create(
        model="whisper-1",  # verify current ID on Models page
        file=f,
        language="en",
        response_format="verbose_json",
    )
print(tx.text)

Preprocessing and splitting

PracticeWhy
Formats listed in docs (mp3, wav, m4a)Fewer codec failures
Split long audio on silence or 5–10 min chunksAvoid size/timeout limits
Mono, light denoiseBetter recognition
Pass language="en" (or target)Reduces mis-detection

Normalize punctuation and filter sensitive terms before downstream use; summarization and classification belong in Responses, not crammed into STT params.

Output formats

  • text: simplest plain text
  • verbose_json: segments with timestamps for player seek
  • srt / vtt: subtitle files—verify player compatibility

TTS: synthesis endpoint

POST /v1/audio/speech converts text to mp3, opus, or pcm.

from pathlib import Path
from openai import OpenAI

client = OpenAI()
speech = client.audio.speech.create(
    model="tts-1",
    voice="nova",
    input="Your ticket has been received. We will reply within two business days.",
    response_format="mp3",
)
Path("reply.mp3").write_bytes(speech.content)
ParameterDev notes
voiceList per docs; pick 1–2 for product consistency
speedIf supported, IVR may use slight speed-up
streamLong read-aloud: generate while playing, lower time-to-first-byte

Caching: fixed phrases (verification codes, standard replies) by (voice, text) hash can cut cost sharply.

Typical pipeline

User recording → object storage → STT → text cleanup → Responses (summary / intent / reply)
                                              ↓
                                    optional TTS → CDN → client playback

Combine with Vision: video audio STT + key-frame vision. Index transcripts via Embeddings guide after PII redaction.

Billing, rate limits, and compliance

  • STT often bills by minute; TTS by characters—see openai.com/api/pricing
  • Exponential backoff on 429 / 5xx; batch via async workers
  • Recording transcription needs user consent and privacy disclosure
  • Medical/legal: human review; don’t TTS impersonate others without authorization

Frequently asked questions

Whisper misses domain jargon?

Set language, post-process with domain glossary in Responses; extreme cases add custom dictionary + spot checks.

File exceeds size limit?

Split on silence or fixed duration; check Batch API in docs.

Can TTS clone a real person?

Follow platform policy; commercial voice cloning needs legal authorization.

Mix REST and Realtime?

Yes: Realtime for live dialog, REST STT for offline archive, shared Responses backend.

Store transcripts directly?

Minimize retention, encrypt, set TTL; redact before DB or index.

Official resources

Next reading

Action path

Today: Transcribe a 1-minute mp3; save verbose_json. Tomorrow: TTS the same reply text; compare tts-1 vs hd tier. This week: Build upload → STT → Responses summary → optional TTS; log STT cost per minute.

Related