Skip to content

Responses API Guide

Last updated:2026-08-12· 17 min read

🚀 Quick access

  • ChatGPT Domestic:Open entry↗
  • Mirror site:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

Responses API Guide

Last updated: 2026-08-12

Overview

If you integrate OpenAI conversational capabilities from scratch today, official docs point to Responses API, not legacy Chat Completions. Responses unifies text/multimodal input, tool use, and conversation state—and is the default home for new features (hosted web search, file search, etc.). Field names, tool list, model IDs, and pricing follow Responses guide and pricing—manage model and schema versions in config.

Positioning: successor to Chat Completions

OpenAI treats Responses as the preferred integration path; Chat Completions is maintenance mode with new capabilities landing on Responses first. Assistants (Thread / Run) converges toward Responses + Agents SDK.

CapabilityChat CompletionsResponses (recommended)
Docs primaryNoYes
Input shapemessages[]instructions + input
Output shapechoices[].messageoutput[] + output_text
MultimodalSupportedSame endpoint—see Vision
Conversation stateClient builds historyprevious_response_id chain
Hosted toolsLimitedweb search, etc. (per docs)

2026 default: new projects use Responses; legacy Completions plans shadow migration with rollback.

Request–response model

POST /v1/responses
  model + instructions + input (+ tools + previous_response_id)
    → output[]: message | function_call | …
    → usage + response.id
FieldDeveloper use
instructionsLong-lived system behavior, format rules
inputString or items (text / image / tool results)
toolsCustom JSON Schema + platform-hosted tools
storeWhen true, later reference by id—mind retention policy
streamSSE for better first-token UX

Minimal runnable example

from openai import OpenAI

client = OpenAI()
resp = client.responses.create(
    model="gpt-4o",  # verify current ID on Models page
    instructions="You are a concise technical assistant. Answer in at most three sentences.",
    input="What is the biggest difference between Responses API and Chat Completions?",
)
print(resp.output_text)
print(resp.usage)
curl https://api.openai.com/v1/responses \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o","instructions":"You are helpful.","input":"Hi"}'

gpt-4o is a placeholder—replace with current Models page ID.

Field mapping (Completions → Responses)

Chat CompletionsResponses
messages[0] role=systeminstructions
messages[-1] role=userinput or input items
choices[0].message.contentoutput_text
tool_calls + manual loopfunction_call in output + return tool result
response_format: json_objectResponses JSON schema / text.format

Side-by-side the same business request in Playground beats hand-translating fields.

Tool use: custom and hosted

Custom functions: declare tools → model returns function_call → server executes → pass result as new input item → create again until final text.

Hosted tools: docs may offer web search, file search, code interpreter—confirm data residency, permissions, and add-on billing (may exceed pure token cost).

Multi-agent orchestration: Agents SDK guide; still calls Responses underneath.

Streaming and idempotency

stream = client.responses.create(
    model="gpt-4o",
    input="Write a four-line poem about API design",
    stream=True,
)
for event in stream:
    pass  # event types per SDK

Side-effect tools (DB writes, email, charges) need idempotency keys so SSE reconnect retries don’t duplicate actions.

Shadow migration in four steps

  1. Pick representative production requests; convert to Responses JSON in Playground
  2. Shadow dual-run: Completions and Responses in parallel—compare quality, tokens, latency
  3. Shift read traffic to Responses; keep Completions rollback switch 2–4 weeks
  4. Subscribe to Changelog for deprecations

Assistants legacy: Assistants guide.

Common stacks

ScenarioStack
RAGEmbeddings retrieval + Responses generation
Receipt OCRVision input items + JSON schema
Long chatprevious_response_id or periodic history summary

Production checklist

  • Keys server-side only; timeout + 429 exponential backoff
  • Log response.id for support
  • Tool allowlist + param validation against prompt injection
  • Monitor tokens + hosted tool surcharges
  • Configure store and data retention per org policy

Frequently asked questions

Will Chat Completions shut down immediately?

Follow official Changelog; usually a transition period with new features on Responses first.

Is output_text enough?

Fine for plain Q&A; tool calls or multi-segment output need iterating output[].

Is Responses more expensive for the same model?

Token unit price usually matches; hosted tools may add fees—compare total bill in shadow.

Can I use REST without SDK?

Yes; fields in API reference; SDKs often track new params faster.

New agent: Responses or Assistants?

New work prefers Responses + Agents SDK; don’t invest in Thread/Run for greenfield.

Official resources

Next reading

Action path

Today: Non-streaming Responses call; print output_text and usage. Tomorrow: Complete one custom function tool loop. This week: Shadow one Completions traffic slice; draft migration timeline.

Related