Skip to content

ChatGPT API Developer Guide

Last updated:2026-08-12· 18 min read

🚀 Quick access

  • ChatGPT Domestic:Open entry↗
  • Mirror site:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

ChatGPT API Developer Guide

Last updated: 2026-08-12

Overview

A working curl is only the start. Production needs key isolation, observability, rate-limit handling, retries, and a migration plan for Responses API / Chat Completions coexisting—Assistants API may be legacy. This guide targets backend and full-stack engineers who’ve read API Quickstart. Fields and endpoints follow API Reference.

[Web / App] ──→ [Your BFF / Gateway] ──→ api.openai.com
                      │
            tenant auth, QPS limits,
            prompt templates, RAG injection,
            logging and token metering
LayerShould doMust not
ClientUI, session idHold OpenAI key
BFFAssemble messages, merge retrievalSend key to browser
WorkerAsync batch, long jobsRaw HTTP without timeout
OpenAIInferenceBusiness RBAC

Auth and key lifecycle

Authorization: Bearer sk-...
Content-Type: application/json
  • Read keys from env vars or Vault; validate presence at startup.
  • Multi-tenant SaaS: don’t give tenants extractable keys—proxy from backend.
  • Rotation: new key live → dual-write verify → switch traffic → revoke old key.

Platform details: Platform Overview.

Choosing an API shape

ShapeEndpoint (example)Best for
Responses API/v1/responsesQuickstart default; tool use
Chat Completions/v1/chat/completionsLegacy, many third-party examples
Assistants API/v1/assistants, etc.Stateful Thread; check legacy status
Embeddings/v1/embeddingsVectors—see embeddings guide

New projects: follow Quickstart.Legacy: set migration milestones; avoid indefinite dual-stack. Assistants migration: Assistants guide.

Request design: messages and roles

{
  "model": "gpt-4o-mini",
  "messages": [
    {"role": "system", "content": "You are order support. Answer shipping and returns only; escalate other topics."},
    {"role": "user", "content": "Where is order 12345?"}
  ],
  "temperature": 0.2,
  "max_tokens": 512
}

gpt-4o-mini is an example model ID—replace with current docs. For the flagship shift, see GPT-6 Astra (gpt-6-astra): keep a rollback to a stable production id and log the verification date.

RolePurpose
systemRules, tool docs, output format
userEnd user or upstream system
assistantPrior replies for multi-turn context
tool / functionTool return (varies by API version)

Multi-turn: your service maintains messages; summarize old turns near context limits. Parameter tuning: Text Generation.

Streaming output (SSE)

Chat UIs benefit from stream: true for lower time-to-first-token:

curl https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "stream": true,
    "messages": [{"role": "user", "content": "Write a five-line poem about morning coding"}]
  }'

Production notes:

  • Cancel upstream when the client disconnects to avoid wasted tokens.
  • Streaming still bills; persist the full concatenated result.
  • Gateway must handle text/event-stream and proxy buffering.

Tool use (Tools / Functions)

{
  "model": "gpt-4o-mini",
  "messages": [{"role": "user", "content": "What's the weather in Shanghai today?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Call when user asks for current weather in a city",
      "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"]
      }
    }
  }]
}
PracticeWhy
Stable schemaEasier prompt regression
Server-side param validationNever execute model-generated SQL/URLs blindly
Timeout and idempotencyGraceful fallback when external APIs fail

Complex orchestration: Agents SDK, Responses API.

Error handling and retries

import time
from openai import OpenAI, RateLimitError, APIStatusError

client = OpenAI()
RETRY = {429, 500, 502, 503, 504}

def chat_with_backoff(**kwargs):
    for i in range(5):
        try:
            return client.chat.completions.create(**kwargs)
        except RateLimitError:
            time.sleep(2 ** i)
        except APIStatusError as e:
            if e.status_code not in RETRY:
                raise
            time.sleep(2 ** i)
    raise RuntimeError("max retries")
  • 401 / 400: config or body errors—don’t blind retry.
  • 429: respect Retry-After when present.
  • Log x-request-id.

Rate limits, cost, and caching

  • Gateway per-tenant QPS so one customer doesn’t exhaust org RPM.
  • Rate-limit long contexts separately; input tokens often dominate cost.
  • Short TTL cache for identical FAQ + system (mind privacy and staleness).
  • Aggregate tokens and error rates by model, route, tenant.

Structured output

  1. Declare JSON schema in system with “JSON only.”
  2. Use official Structured Outputs / JSON mode when the model supports it.
  3. Validate with JSON Schema server-side; retry or degrade on failure.

Prompt templates: Prompt Engineering.

Production launch checklist

Security

  • Keys not in frontend or mobile
  • Redacted input/output logs
  • Human review for high-risk scenarios

Reliability

  • Timeouts (connect + read)
  • 429/5xx backoff
  • Circuit breaker on sustained failures

Quality

  • Golden prompts before release
  • Versioned system prompts (Git / CMS)

Compliance

  • User consent and data retention policy
  • Content safety filters per business rules

Relationship to Assistants API

Assistants offers Assistant / Thread / Run with File Search built in. Official guidance may steer toward Responses API or Agents SDK. Before new features, read migration notes in Assistants guide—don’t assume Assistants is the long-term default.

Frequently asked questions

Can multiple microservices share one key?

Possible but not ideal; at minimum separate by environment, ideally by service for leak isolation and quota analysis.

How do I implement multi-turn “memory”?

Self-manage messages or session summaries; or use stateful APIs (Assistants, etc.) and weigh vendor lock-in and cost.

Is stream vs non-stream quality different?

Same model and parameters yield the same distribution; difference is UX and implementation, not quality.

BYOK (bring your own key)?

Feasible with strict isolation and audit; most products proxy from backend.

How do I compare OpenAI vs Claude cost?

Run the same eval set; log tokens and latency; check OpenAI pricing and Anthropic pricing.

Official resources

Next reading

Action path

Today: Add timeout and 429 backoff to existing calls. Tomorrow: Enable streaming and measure time-to-first-token. This week: Build a 20-prompt golden set and a one-page gateway architecture doc for team review.

Related