Skip to content

Assistants API

Last updated:2026-08-12· 17 min read

🚀 Quick access

  • ChatGPT Domestic:Open entry↗
  • Mirror site:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

Assistants API

Last updated: 2026-08-12

Overview

The Assistants API models stateful conversation with Assistant + Thread + Run, with built-in Code Interpreter and File Search—good for quick “file-backed Q&A bot” POCs. Platform evolution may favor Responses API and Agents SDK for new work; Assistants may be in maintenance or migration—read official docs deprecation and migration before committing.

Important: If Assistants is marked legacy, don’t default for new projects; plan migration for existing ones. Endpoints and objects follow API reference; model below is example only.

Core object model

Assistant (config: instructions / model / tools)
    │
Thread (conversation container)
    └── Messages
         │
Run (one execution: queued → in_progress → completed)
ObjectRoleYou typically persist
AssistantPersona and tool configassistant_id
ThreadMulti-turn historyPer-user thread_id
MessageSingle input/outputOptional local mirror
RunOne model executionPoll / webhook
File / Vector StoreRetrieval assetsFile id list

Vs stateless APIs: Chat Completions / Responses—you maintain messages[]; Assistants stores Thread on OpenAI—convenient but vendor-bound with different billing dimensions.

New vs legacy: how to choose

SituationRecommendation
Quickstart points to ResponsesRead Responses API Guide; don’t new-ship Assistants
Production Threads + Vector StoresMaintain + plan migration
POC for file Q&AAssistants is fast to start; reassess after POC
Multi-tenant, fine RAG controlPrefer embeddings + Completions / Responses

Three-step curl flow (shape example)

1. Create Assistant

curl https://api.openai.com/v1/assistants \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -H "OpenAI-Beta: assistants=v2" \
  -d '{
    "name": "policy-bot",
    "instructions": "Answer only from retrieved files; say not found if insufficient.",
    "model": "gpt-4o-mini",
    "tools": [{"type": "file_search"}]
  }'

2. Create Thread and add message

curl https://api.openai.com/v1/threads \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -H "OpenAI-Beta: assistants=v2" \
  -d '{
    "messages": [{"role": "user", "content": "What is the return window in the handbook?"}]
  }'

3. Create Run

curl https://api.openai.com/v1/threads/thread_abc/runs \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -H "OpenAI-Beta: assistants=v2" \
  -d '{"assistant_id": "asst_abc"}'

Production should:

  • Poll GET /threads/{id}/runs/{run_id} or use streaming / webhook
  • Handle requires_action (submit function output)
  • Set Run timeout and failure retry

File Search workflow

  1. POST /v1/files upload (purpose per docs)
  2. Create Vector Store and attach to Assistant
  3. User question → Run → model retrieves via file_search
PracticeNotes
File versionsReindex on update; clean old vectors
Multi-tenantDon’t share one Assistant for confidential files
EvalCases where doc has answer vs doesn’t

Self-built RAG: Embeddings guide + Chat Completions / Responses.

Code Interpreter cautions

  • Good for data analysis POC; assess sandbox security, cost, reproducibility for production.
  • Don’t pass secrets or PII.
  • If output files are downloadable, limit size and scan.

Migration to Responses / Agents SDK

CapabilityAssistantsResponses / Agents
Stateful sessionBuilt-in ThreadSelf-managed session or SDK
ToolsBuilt-in + functionsUnified tools (per docs)
Long-term supportMay weakenQuickstart direction

Generic migration steps (follow official migration guide):

  1. instructions → system template (Prompt Engineering)
  2. file_search → self vector store or new official option
  3. Thread history → import to DB or cold-start with summary
  4. Dual-write eval, then cut traffic

Billing and monitoring

  • Runs consume model tokens; file_search, code_interpreter may add charges.
  • Polling doesn’t use tokens but consumes QPS.
  • Aggregate failure rate by assistant_id and Run status.

Production checklist (if you continue)

  • Confirm API version not sunset (e.g. OpenAI-Beta: assistants=v2)
  • Tenant-level Assistant / Vector Store isolation
  • Run timeout and user-facing fallback copy
  • Test coverage for requires_action
  • Migration plan linked to Responses API

Frequently asked questions

Can new projects still choose Assistants?

Check official docs first. If legacy or Responses is recommended, don’t default to Assistants.

How long is Thread data kept?

Per OpenAI data policy; compliant businesses should back up critical conversations.

Can this replace “free ChatGPT” backend?

No. API bills per usage; Runs + tools can cost more; ChatGPT subscription doesn’t cover API.

file_search vs self-built RAG?

Assistants is faster to start; self RAG gives better multi-tenant and cost control.

Run stuck in in_progress?

Long tool runs or queue; set timeout, check status page, cancel and retry if needed.

Official resources

Next reading

Action path

Today: Confirm Assistants status (active / legacy / migration) in official docs. Tomorrow: Inventory all assistant_id and file dependencies for legacy systems. This week: Build a 5-question migration eval set and estimate cost against Responses Quickstart.

Related