Skip to content

Embeddings

Last updated:2026-08-12· 15 min read

🚀 Quick access

  • ChatGPT Domestic:Open entry↗
  • Mirror site:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

Embeddings

Last updated: 2026-08-12

Overview

The Embeddings API maps text to high-dimensional float vectors so semantically similar sentences sit closer in vector space. It underpins semantic search, deduplication, recommendation recall, and RAG (retrieval-augmented generation). This guide covers endpoints, indexing pipelines, evaluation, and production pitfalls for backend and data engineers. Model IDs, dimensions, and token pricing follow OpenAI docs and pricing—manage model names in config, not hardcoded strings.

DimensionKeyword / BM25Embeddings
Match logicLiteral overlapSemantic similarity
“Refund” finding “cancel subscription”Often missesUsually recalls
Exact SKU, order idStrongWeak—keep SQL / ES
Cross-language near-synonymsNeeds synonym listsRelatively friendly
Final answer generationCannotCannot—needs Responses API

Division of labor: Embeddings find chunks; generation models write answers.

API endpoint and call patterns

Standard request: POST /v1/embeddings. input accepts a string or array; batching reduces HTTP overhead within per-request token limits (see API Reference).

curl https://api.openai.com/v1/embeddings \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "text-embedding-3-small",
    "input": ["Return policy summary", "How to cancel subscription"],
    "encoding_format": "float"
  }'
from openai import OpenAI

client = OpenAI()
resp = client.embeddings.create(
    model="text-embedding-3-small",  # verify current ID on Models page
    input="OpenAI embeddings bill by input tokens",
)
print(len(resp.data[0].embedding), resp.usage.total_tokens)

Production habits:

  • Cache vectors when document content is unchanged (DB / Redis / object storage)
  • Large indexing: queue + rate limit; exponential backoff on 429
  • Log usage.total_tokens for cost allocation

Model selection and version lock

ConsiderationApproach
CostBaseline Recall with small tier; compare large if needed
MultilingualEval with real queries in target languages
StorageSome models support dimensions reduction—requires full re-embed
UpgradeNew model = rebuild index + regression

Rule: index and online queries must use the same embedding model; mixing breaks recall.

RAG indexing pipeline

Raw docs → clean → chunk → batch embed → write vector store (with metadata)
User query → embed query → filter + ANN → Top-K → assemble context → Responses generation

Chunking strategies

StrategyBest forNotes
Fixed token + overlapGeneral wiki512–1024 tokens, 10–20% overlap start
By heading / sectionTech manualsRe-chunk long sections
Parent-childLong reportsSmall chunks retrieve, large chunks generate
Sentence-levelFAQEasy to lose context

Four ways to improve recall

  1. Hybrid retrieval: vector Top-K + BM25, RRF fusion
  2. Metadata pre-filter: version, product line, permission tags first
  3. Query rewrite: light model turns colloquial questions retrieval-friendly
  4. Rerank: rerank Top-20 to Top-5 (higher latency and cost)

Generation constraints: Prompt Engineering guide.

Vector stores and similarity

The API returns float arrays; persistence and ANN are your choice: pgvector, Pinecone, Milvus, Elasticsearch dense_vector, etc. OpenAI vectors are often L2-normalized—cosine similarity equals dot product.

Choose by: data volume, QPS, hybrid filters, ops capacity, multi-tenant isolation.

Billing, rate limits, and security

  • Billing: input tokens; full index cost ≈ total doc tokens × unit price; each query also embeds (cache hot queries)
  • Rate limits: backoff on 429; offline indexing via worker pools
  • Security: redact PII before indexing; permission-check retrieval before prompt injection; read API data retention policy

Pricing: openai.com/api/pricing

Copy-ready RAG context template

[System]
You are an enterprise knowledge assistant. Answer only from provided context.
If context is insufficient, say "Not mentioned in the materials" and list what's missing.
Do not invent links or legal citations.

[User]
context:
"""
[paste retrieved redacted passages with doc_id]
"""
Question: [user question]

Frequently asked questions

How do Embeddings and fine-tuning work together?

Embeddings handle retrieval; fine-tuning handles generation style and format. Common RAG stack: Embeddings + base model without fine-tuning. See Fine-tuning guide.

Retrieval is correct but answers hallucinate?

Often missing “answer only from context,” noisy oversized chunks, or no citation requirement. Tighten instructions and validate server-side.

Can I embed images directly?

Text Embeddings API is for strings; image understanding uses Vision guide, or OCR → text → embed.

Worth switching to open-source embeddings?

Open-source enables local deploy; OpenAI hosted saves ops. Decide with business query set on Recall@K and P95 latency.

After dimension reduction, rebuild index?

Yes—re-embed entire corpus, rebuild ANN, then A/B.

Official resources

Next reading

Action path

Today: Embed 3 similar and 3 unrelated sentence pairs; compute cosine to feel semantic distance. Tomorrow: Chunk one Markdown doc at 800 tokens + 15% overlap; batch write to pgvector. This week: 30 labeled Q&A pairs for Recall@5; wire Responses for minimal RAG loop.

Related