Skip to content

Fine-tuning

Last updated:2026-08-12· 16 min read

🚀 Quick access

  • ChatGPT Domestic:Open entry↗
  • Mirror site:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

Fine-tuning

Last updated: 2026-08-12

Overview

Fine-tuning continues training a base model on your high-quality examples so outputs stabilize on fixed format, tone, classification boundaries, or domain terms. OpenAI’s hosted flow: upload JSONL → create fine-tuning job → get ft:... model ID. Eligible base models, data format, training and inference pricing follow Fine-tuning docs and pricing—verify before each run.

Do you actually need fine-tuning?

Need latest external knowledge? ──yes──→ prefer RAG (Embeddings + Responses)
       │
       no
       ↓
Prompt + JSON Schema already stable? ──yes──→ skip fine-tuning
       │
       no
       ↓
Large labeled set, stable distribution, prompt exhausted? ──yes──→ consider fine-tuning
PathBetter when
Prompt + structured outputFormat controllable; examples fit in context
RAGKnowledge updates often; citations needed
Tools / Responses APIQuery DB, HTTP, run code
Fine-tuningHuge stable samples; style / classification / JSON compliance still fail
Agents SDKMulti-step orchestration—not weight training

Fine-tuning does not inject new facts automatically; factual errors still need RAG or tools.

Training data: JSONL spec

Supervised fine-tuning typically uses JSONL, one row per example. Chat models often use a messages array:

{"messages": [
  {"role": "system", "content": "You are a ticket classifier. Output one JSON line, no markdown."},
  {"role": "user", "content": "Package hasn't arrived in three days, I want to complain"},
  {"role": "assistant", "content": "{\"intent\":\"logistics\",\"priority\":\"high\"}"}
]}

Data quality checklist

  • Complex tasks often need 500+ quality samples; smoke-train 50 rows first to validate format
  • assistant is the gold answer—consistent format, no extra chitchat
  • Cover refusal, ambiguity, edge cases
  • 90/10 train/validation split; avoid user-level leakage
  • Remove PII; comply with copyright and API data policy
  • No undocumented fields or wrong role order (common failure causes)

Training job workflow

StepActionNotes
1. UploadPOST /v1/files (purpose: fine-tune)Size/row limits in docs
2. CreatePOST /v1/fine_tuning/jobsbase_model must be eligible
3. MonitorPoll statusOn failed, download error file
4. Callmodel="ft:..."Same as base model invocation
from openai import OpenAI

client = OpenAI()
job = client.fine_tuning.jobs.create(
    training_file="file-abc123",
    model="gpt-4o-mini-2024-07-18",  # per eligible list in docs
    validation_file="file-def456",
)
print(job.id, job.status)

Hyperparameters (epochs, learning_rate_multiplier, etc.): start from doc recommendations; grid search on small validation set.

Evaluation: beyond training loss

DimensionApproach
Baseline comparisonHeld-out set vs best prompt + base model
Format validityJSON parse, schema pass rate
Business metricsF1, accuracy, human satisfaction
Safety regressionHarmful requests must still refuse
Online A/B5% canary—latency, cost, error rate

Lower loss ≠ better business outcomes—trust held-out business metrics.

Launch, versioning, and cost

  • Version snapshot: keep ft:... ID, training set hash, hyperparams, eval report
  • Coexist strategy: canary new ft; roll back to base model if metrics drop
  • Cost: training (tokens × epochs) + inference (often above same-tier base)—estimate ROI on pricing page
  • Retrain triggers: rule changes, new intent share threshold, eval decline
  • Safety: fine-tuning doesn’t replace moderation; review sensitive outputs

Combining with other capabilities

User input → RAG retrieve context → ft model outputs JSON → server schema validate → persist
Multi-step → Agents SDK orchestration → call ft model only at classification / formatting nodes

Frequently asked questions

Is 100 examples enough?

Maybe for simple binary classification but overfitting risk is high. Always hold out validation; stop if worse than prompt baseline.

How do I call a fine-tuned model?

Set model to full ft:... ID in Chat / Responses; SDK usage matches base models.

Can I fine-tune embedding models?

Check current product; most RAG teams tune chunking, hybrid search, and rerank first.

Job failed—how to debug?

Download error file; common issues: JSONL format, row too long, invalid role, empty assistant.

Will OpenAI train on my data?

Read latest API data usage and enterprise agreements.

Official resources

Next reading

Action path

Today: Collect 5 samples where prompt tuning still fails; separate knowledge gaps vs format/style issues. Tomorrow: If latter, hand-write 20 JSONL lines and submit smoke job. This week: On held-out 50, compare base vs ft format pass rate before scaling data and launch.

Related