Fine-tuning
Last updated:2026-08-12· 16 min read
🚀 Quick access
- ChatGPT Domestic:Open entry↗
- Mirror site:Open mirror↗
- Official ChatGPT:chatgpt.com ↗

Last updated: 2026-08-12
Overview
Fine-tuning continues training a base model on your high-quality examples so outputs stabilize on fixed format, tone, classification boundaries, or domain terms. OpenAI’s hosted flow: upload JSONL → create fine-tuning job → get ft:... model ID. Eligible base models, data format, training and inference pricing follow Fine-tuning docs and pricing—verify before each run.
Do you actually need fine-tuning?
Need latest external knowledge? ──yes──→ prefer RAG (Embeddings + Responses)
│
no
↓
Prompt + JSON Schema already stable? ──yes──→ skip fine-tuning
│
no
↓
Large labeled set, stable distribution, prompt exhausted? ──yes──→ consider fine-tuning
| Path | Better when |
|---|---|
| Prompt + structured output | Format controllable; examples fit in context |
| RAG | Knowledge updates often; citations needed |
| Tools / Responses API | Query DB, HTTP, run code |
| Fine-tuning | Huge stable samples; style / classification / JSON compliance still fail |
| Agents SDK | Multi-step orchestration—not weight training |
Fine-tuning does not inject new facts automatically; factual errors still need RAG or tools.
Training data: JSONL spec
Supervised fine-tuning typically uses JSONL, one row per example. Chat models often use a messages array:
{"messages": [
{"role": "system", "content": "You are a ticket classifier. Output one JSON line, no markdown."},
{"role": "user", "content": "Package hasn't arrived in three days, I want to complain"},
{"role": "assistant", "content": "{\"intent\":\"logistics\",\"priority\":\"high\"}"}
]}
Data quality checklist
- Complex tasks often need 500+ quality samples; smoke-train 50 rows first to validate format
-
assistantis the gold answer—consistent format, no extra chitchat - Cover refusal, ambiguity, edge cases
- 90/10 train/validation split; avoid user-level leakage
- Remove PII; comply with copyright and API data policy
- No undocumented fields or wrong role order (common failure causes)
Training job workflow
| Step | Action | Notes |
|---|---|---|
| 1. Upload | POST /v1/files (purpose: fine-tune) | Size/row limits in docs |
| 2. Create | POST /v1/fine_tuning/jobs | base_model must be eligible |
| 3. Monitor | Poll status | On failed, download error file |
| 4. Call | model="ft:..." | Same as base model invocation |
from openai import OpenAI
client = OpenAI()
job = client.fine_tuning.jobs.create(
training_file="file-abc123",
model="gpt-4o-mini-2024-07-18", # per eligible list in docs
validation_file="file-def456",
)
print(job.id, job.status)
Hyperparameters (epochs, learning_rate_multiplier, etc.): start from doc recommendations; grid search on small validation set.
Evaluation: beyond training loss
| Dimension | Approach |
|---|---|
| Baseline comparison | Held-out set vs best prompt + base model |
| Format validity | JSON parse, schema pass rate |
| Business metrics | F1, accuracy, human satisfaction |
| Safety regression | Harmful requests must still refuse |
| Online A/B | 5% canary—latency, cost, error rate |
Lower loss ≠ better business outcomes—trust held-out business metrics.
Launch, versioning, and cost
- Version snapshot: keep
ft:...ID, training set hash, hyperparams, eval report - Coexist strategy: canary new ft; roll back to base model if metrics drop
- Cost: training (tokens × epochs) + inference (often above same-tier base)—estimate ROI on pricing page
- Retrain triggers: rule changes, new intent share threshold, eval decline
- Safety: fine-tuning doesn’t replace moderation; review sensitive outputs
Combining with other capabilities
User input → RAG retrieve context → ft model outputs JSON → server schema validate → persist
Multi-step → Agents SDK orchestration → call ft model only at classification / formatting nodes
Frequently asked questions
Is 100 examples enough?
Maybe for simple binary classification but overfitting risk is high. Always hold out validation; stop if worse than prompt baseline.
How do I call a fine-tuned model?
Set model to full ft:... ID in Chat / Responses; SDK usage matches base models.
Can I fine-tune embedding models?
Check current product; most RAG teams tune chunking, hybrid search, and rerank first.
Job failed—how to debug?
Download error file; common issues: JSONL format, row too long, invalid role, empty assistant.
Will OpenAI train on my data?
Read latest API data usage and enterprise agreements.
Official resources
Next reading
Action path
Today: Collect 5 samples where prompt tuning still fails; separate knowledge gaps vs format/style issues. Tomorrow: If latter, hand-write 20 JSONL lines and submit smoke job. This week: On held-out 50, compare base vs ft format pass rate before scaling data and launch.
Related
OpenAI Dev Overview
2026 OpenAI developer map: how ChatGPT web, Platform console, and APIs divide work—and the reading order from first call to production.
OpenAI Platform Overview
platform.openai.com console, doc navigation, Playground, usage billing, and org management—how developers find API information efficiently.
OpenAI API Quickstart
From Platform account and API key to your first OpenAI call: Responses/Completions examples, billing, rate limits, and a security checklist (2026 hands-on).
ChatGPT API Developer Guide
Production OpenAI API integration: architecture, auth, streaming, tool use, rate-limit retries, and a launch checklist.