Embeddings
Last updated:2026-08-12· 15 min read
🚀 Quick access
- ChatGPT Domestic:Open entry↗
- Mirror site:Open mirror↗
- Official ChatGPT:chatgpt.com ↗

Last updated: 2026-08-12
Overview
The Embeddings API maps text to high-dimensional float vectors so semantically similar sentences sit closer in vector space. It underpins semantic search, deduplication, recommendation recall, and RAG (retrieval-augmented generation). This guide covers endpoints, indexing pipelines, evaluation, and production pitfalls for backend and data engineers. Model IDs, dimensions, and token pricing follow OpenAI docs and pricing—manage model names in config, not hardcoded strings.
Embeddings vs keyword search
| Dimension | Keyword / BM25 | Embeddings |
|---|---|---|
| Match logic | Literal overlap | Semantic similarity |
| “Refund” finding “cancel subscription” | Often misses | Usually recalls |
| Exact SKU, order id | Strong | Weak—keep SQL / ES |
| Cross-language near-synonyms | Needs synonym lists | Relatively friendly |
| Final answer generation | Cannot | Cannot—needs Responses API |
Division of labor: Embeddings find chunks; generation models write answers.
API endpoint and call patterns
Standard request: POST /v1/embeddings. input accepts a string or array; batching reduces HTTP overhead within per-request token limits (see API Reference).
curl https://api.openai.com/v1/embeddings \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "text-embedding-3-small",
"input": ["Return policy summary", "How to cancel subscription"],
"encoding_format": "float"
}'
from openai import OpenAI
client = OpenAI()
resp = client.embeddings.create(
model="text-embedding-3-small", # verify current ID on Models page
input="OpenAI embeddings bill by input tokens",
)
print(len(resp.data[0].embedding), resp.usage.total_tokens)
Production habits:
- Cache vectors when document content is unchanged (DB / Redis / object storage)
- Large indexing: queue + rate limit; exponential backoff on 429
- Log
usage.total_tokensfor cost allocation
Model selection and version lock
| Consideration | Approach |
|---|---|
| Cost | Baseline Recall with small tier; compare large if needed |
| Multilingual | Eval with real queries in target languages |
| Storage | Some models support dimensions reduction—requires full re-embed |
| Upgrade | New model = rebuild index + regression |
Rule: index and online queries must use the same embedding model; mixing breaks recall.
RAG indexing pipeline
Raw docs → clean → chunk → batch embed → write vector store (with metadata)
User query → embed query → filter + ANN → Top-K → assemble context → Responses generation
Chunking strategies
| Strategy | Best for | Notes |
|---|---|---|
| Fixed token + overlap | General wiki | 512–1024 tokens, 10–20% overlap start |
| By heading / section | Tech manuals | Re-chunk long sections |
| Parent-child | Long reports | Small chunks retrieve, large chunks generate |
| Sentence-level | FAQ | Easy to lose context |
Four ways to improve recall
- Hybrid retrieval: vector Top-K + BM25, RRF fusion
- Metadata pre-filter: version, product line, permission tags first
- Query rewrite: light model turns colloquial questions retrieval-friendly
- Rerank: rerank Top-20 to Top-5 (higher latency and cost)
Generation constraints: Prompt Engineering guide.
Vector stores and similarity
The API returns float arrays; persistence and ANN are your choice: pgvector, Pinecone, Milvus, Elasticsearch dense_vector, etc. OpenAI vectors are often L2-normalized—cosine similarity equals dot product.
Choose by: data volume, QPS, hybrid filters, ops capacity, multi-tenant isolation.
Billing, rate limits, and security
- Billing: input tokens; full index cost ≈ total doc tokens × unit price; each query also embeds (cache hot queries)
- Rate limits: backoff on 429; offline indexing via worker pools
- Security: redact PII before indexing; permission-check retrieval before prompt injection; read API data retention policy
Pricing: openai.com/api/pricing
Copy-ready RAG context template
[System]
You are an enterprise knowledge assistant. Answer only from provided context.
If context is insufficient, say "Not mentioned in the materials" and list what's missing.
Do not invent links or legal citations.
[User]
context:
"""
[paste retrieved redacted passages with doc_id]
"""
Question: [user question]
Frequently asked questions
How do Embeddings and fine-tuning work together?
Embeddings handle retrieval; fine-tuning handles generation style and format. Common RAG stack: Embeddings + base model without fine-tuning. See Fine-tuning guide.
Retrieval is correct but answers hallucinate?
Often missing “answer only from context,” noisy oversized chunks, or no citation requirement. Tighten instructions and validate server-side.
Can I embed images directly?
Text Embeddings API is for strings; image understanding uses Vision guide, or OCR → text → embed.
Worth switching to open-source embeddings?
Open-source enables local deploy; OpenAI hosted saves ops. Decide with business query set on Recall@K and P95 latency.
After dimension reduction, rebuild index?
Yes—re-embed entire corpus, rebuild ANN, then A/B.
Official resources
Next reading
- Responses API Guide (preferred RAG generation side)
- Fine-tuning
- Prompt Engineering
- OpenAI API Quickstart
Action path
Today: Embed 3 similar and 3 unrelated sentence pairs; compute cosine to feel semantic distance. Tomorrow: Chunk one Markdown doc at 800 tokens + 15% overlap; batch write to pgvector. This week: 30 labeled Q&A pairs for Recall@5; wire Responses for minimal RAG loop.
Related
OpenAI Dev Overview
2026 OpenAI developer map: how ChatGPT web, Platform console, and APIs divide work—and the reading order from first call to production.
OpenAI Platform Overview
platform.openai.com console, doc navigation, Playground, usage billing, and org management—how developers find API information efficiently.
OpenAI API Quickstart
From Platform account and API key to your first OpenAI call: Responses/Completions examples, billing, rate limits, and a security checklist (2026 hands-on).
ChatGPT API Developer Guide
Production OpenAI API integration: architecture, auth, streaming, tool use, rate-limit retries, and a launch checklist.