Responses API Guide
Last updated:2026-08-12· 17 min read
🚀 Quick access
- ChatGPT Domestic:Open entry↗
- Mirror site:Open mirror↗
- Official ChatGPT:chatgpt.com ↗

Last updated: 2026-08-12
Overview
If you integrate OpenAI conversational capabilities from scratch today, official docs point to Responses API, not legacy Chat Completions. Responses unifies text/multimodal input, tool use, and conversation state—and is the default home for new features (hosted web search, file search, etc.). Field names, tool list, model IDs, and pricing follow Responses guide and pricing—manage model and schema versions in config.
Positioning: successor to Chat Completions
OpenAI treats Responses as the preferred integration path; Chat Completions is maintenance mode with new capabilities landing on Responses first. Assistants (Thread / Run) converges toward Responses + Agents SDK.
| Capability | Chat Completions | Responses (recommended) |
|---|---|---|
| Docs primary | No | Yes |
| Input shape | messages[] | instructions + input |
| Output shape | choices[].message | output[] + output_text |
| Multimodal | Supported | Same endpoint—see Vision |
| Conversation state | Client builds history | previous_response_id chain |
| Hosted tools | Limited | web search, etc. (per docs) |
2026 default: new projects use Responses; legacy Completions plans shadow migration with rollback.
Request–response model
POST /v1/responses
model + instructions + input (+ tools + previous_response_id)
→ output[]: message | function_call | …
→ usage + response.id
| Field | Developer use |
|---|---|
instructions | Long-lived system behavior, format rules |
input | String or items (text / image / tool results) |
tools | Custom JSON Schema + platform-hosted tools |
store | When true, later reference by id—mind retention policy |
stream | SSE for better first-token UX |
Minimal runnable example
from openai import OpenAI
client = OpenAI()
resp = client.responses.create(
model="gpt-4o", # verify current ID on Models page
instructions="You are a concise technical assistant. Answer in at most three sentences.",
input="What is the biggest difference between Responses API and Chat Completions?",
)
print(resp.output_text)
print(resp.usage)
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-4o","instructions":"You are helpful.","input":"Hi"}'
gpt-4o is a placeholder—replace with current Models page ID.
Field mapping (Completions → Responses)
| Chat Completions | Responses |
|---|---|
messages[0] role=system | instructions |
messages[-1] role=user | input or input items |
choices[0].message.content | output_text |
tool_calls + manual loop | function_call in output + return tool result |
response_format: json_object | Responses JSON schema / text.format |
Side-by-side the same business request in Playground beats hand-translating fields.
Tool use: custom and hosted
Custom functions: declare tools → model returns function_call → server executes → pass result as new input item → create again until final text.
Hosted tools: docs may offer web search, file search, code interpreter—confirm data residency, permissions, and add-on billing (may exceed pure token cost).
Multi-agent orchestration: Agents SDK guide; still calls Responses underneath.
Streaming and idempotency
stream = client.responses.create(
model="gpt-4o",
input="Write a four-line poem about API design",
stream=True,
)
for event in stream:
pass # event types per SDK
Side-effect tools (DB writes, email, charges) need idempotency keys so SSE reconnect retries don’t duplicate actions.
Shadow migration in four steps
- Pick representative production requests; convert to Responses JSON in Playground
- Shadow dual-run: Completions and Responses in parallel—compare quality, tokens, latency
- Shift read traffic to Responses; keep Completions rollback switch 2–4 weeks
- Subscribe to Changelog for deprecations
Assistants legacy: Assistants guide.
Common stacks
| Scenario | Stack |
|---|---|
| RAG | Embeddings retrieval + Responses generation |
| Receipt OCR | Vision input items + JSON schema |
| Long chat | previous_response_id or periodic history summary |
Production checklist
- Keys server-side only; timeout + 429 exponential backoff
- Log
response.idfor support - Tool allowlist + param validation against prompt injection
- Monitor tokens + hosted tool surcharges
- Configure
storeand data retention per org policy
Frequently asked questions
Will Chat Completions shut down immediately?
Follow official Changelog; usually a transition period with new features on Responses first.
Is output_text enough?
Fine for plain Q&A; tool calls or multi-segment output need iterating output[].
Is Responses more expensive for the same model?
Token unit price usually matches; hosted tools may add fees—compare total bill in shadow.
Can I use REST without SDK?
Yes; fields in API reference; SDKs often track new params faster.
New agent: Responses or Assistants?
New work prefers Responses + Agents SDK; don’t invest in Thread/Run for greenfield.
Official resources
Next reading
Action path
Today: Non-streaming Responses call; print output_text and usage. Tomorrow: Complete one custom function tool loop. This week: Shadow one Completions traffic slice; draft migration timeline.
Related
OpenAI Dev Overview
2026 OpenAI developer map: how ChatGPT web, Platform console, and APIs divide work—and the reading order from first call to production.
OpenAI Platform Overview
platform.openai.com console, doc navigation, Playground, usage billing, and org management—how developers find API information efficiently.
OpenAI API Quickstart
From Platform account and API key to your first OpenAI call: Responses/Completions examples, billing, rate limits, and a security checklist (2026 hands-on).
ChatGPT API Developer Guide
Production OpenAI API integration: architecture, auth, streaming, tool use, rate-limit retries, and a launch checklist.