Assistants API
Last updated:2026-08-12· 17 min read
🚀 Quick access
- ChatGPT Domestic:Open entry↗
- Mirror site:Open mirror↗
- Official ChatGPT:chatgpt.com ↗

Last updated: 2026-08-12
Overview
The Assistants API models stateful conversation with Assistant + Thread + Run, with built-in Code Interpreter and File Search—good for quick “file-backed Q&A bot” POCs. Platform evolution may favor Responses API and Agents SDK for new work; Assistants may be in maintenance or migration—read official docs deprecation and migration before committing.
Important: If Assistants is marked legacy, don’t default for new projects; plan migration for existing ones. Endpoints and objects follow API reference;
modelbelow is example only.
Core object model
Assistant (config: instructions / model / tools)
│
Thread (conversation container)
└── Messages
│
Run (one execution: queued → in_progress → completed)
| Object | Role | You typically persist |
|---|---|---|
| Assistant | Persona and tool config | assistant_id |
| Thread | Multi-turn history | Per-user thread_id |
| Message | Single input/output | Optional local mirror |
| Run | One model execution | Poll / webhook |
| File / Vector Store | Retrieval assets | File id list |
Vs stateless APIs: Chat Completions / Responses—you maintain messages[]; Assistants stores Thread on OpenAI—convenient but vendor-bound with different billing dimensions.
New vs legacy: how to choose
| Situation | Recommendation |
|---|---|
| Quickstart points to Responses | Read Responses API Guide; don’t new-ship Assistants |
| Production Threads + Vector Stores | Maintain + plan migration |
| POC for file Q&A | Assistants is fast to start; reassess after POC |
| Multi-tenant, fine RAG control | Prefer embeddings + Completions / Responses |
Three-step curl flow (shape example)
1. Create Assistant
curl https://api.openai.com/v1/assistants \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-H "OpenAI-Beta: assistants=v2" \
-d '{
"name": "policy-bot",
"instructions": "Answer only from retrieved files; say not found if insufficient.",
"model": "gpt-4o-mini",
"tools": [{"type": "file_search"}]
}'
2. Create Thread and add message
curl https://api.openai.com/v1/threads \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-H "OpenAI-Beta: assistants=v2" \
-d '{
"messages": [{"role": "user", "content": "What is the return window in the handbook?"}]
}'
3. Create Run
curl https://api.openai.com/v1/threads/thread_abc/runs \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-H "OpenAI-Beta: assistants=v2" \
-d '{"assistant_id": "asst_abc"}'
Production should:
- Poll
GET /threads/{id}/runs/{run_id}or use streaming / webhook - Handle
requires_action(submit function output) - Set Run timeout and failure retry
File Search workflow
POST /v1/filesupload (purpose per docs)- Create Vector Store and attach to Assistant
- User question → Run → model retrieves via file_search
| Practice | Notes |
|---|---|
| File versions | Reindex on update; clean old vectors |
| Multi-tenant | Don’t share one Assistant for confidential files |
| Eval | Cases where doc has answer vs doesn’t |
Self-built RAG: Embeddings guide + Chat Completions / Responses.
Code Interpreter cautions
- Good for data analysis POC; assess sandbox security, cost, reproducibility for production.
- Don’t pass secrets or PII.
- If output files are downloadable, limit size and scan.
Migration to Responses / Agents SDK
| Capability | Assistants | Responses / Agents |
|---|---|---|
| Stateful session | Built-in Thread | Self-managed session or SDK |
| Tools | Built-in + functions | Unified tools (per docs) |
| Long-term support | May weaken | Quickstart direction |
Generic migration steps (follow official migration guide):
instructions→ system template (Prompt Engineering)- file_search → self vector store or new official option
- Thread history → import to DB or cold-start with summary
- Dual-write eval, then cut traffic
Billing and monitoring
- Runs consume model tokens; file_search, code_interpreter may add charges.
- Polling doesn’t use tokens but consumes QPS.
- Aggregate failure rate by
assistant_idand Run status.
Production checklist (if you continue)
- Confirm API version not sunset (e.g.
OpenAI-Beta: assistants=v2) - Tenant-level Assistant / Vector Store isolation
- Run timeout and user-facing fallback copy
- Test coverage for
requires_action - Migration plan linked to Responses API
Frequently asked questions
Can new projects still choose Assistants?
Check official docs first. If legacy or Responses is recommended, don’t default to Assistants.
How long is Thread data kept?
Per OpenAI data policy; compliant businesses should back up critical conversations.
Can this replace “free ChatGPT” backend?
No. API bills per usage; Runs + tools can cost more; ChatGPT subscription doesn’t cover API.
file_search vs self-built RAG?
Assistants is faster to start; self RAG gives better multi-tenant and cost control.
Run stuck in in_progress?
Long tool runs or queue; set timeout, check status page, cancel and retry if needed.
Official resources
- Assistants API docs (search Assistants)
- Responses API docs
- OpenAI Platform
- API pricing
Next reading
Action path
Today: Confirm Assistants status (active / legacy / migration) in official docs. Tomorrow: Inventory all assistant_id and file dependencies for legacy systems. This week: Build a 5-question migration eval set and estimate cost against Responses Quickstart.
Related
OpenAI Dev Overview
2026 OpenAI developer map: how ChatGPT web, Platform console, and APIs divide work—and the reading order from first call to production.
OpenAI Platform Overview
platform.openai.com console, doc navigation, Playground, usage billing, and org management—how developers find API information efficiently.
OpenAI API Quickstart
From Platform account and API key to your first OpenAI call: Responses/Completions examples, billing, rate limits, and a security checklist (2026 hands-on).
ChatGPT API Developer Guide
Production OpenAI API integration: architecture, auth, streaming, tool use, rate-limit retries, and a launch checklist.