Skip to content

Text Generation

Last updated:2026-08-12· 17 min read

🚀 Quick access

  • ChatGPT Domestic:Open entry↗
  • Mirror site:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

Text Generation

Last updated: 2026-08-12

Overview

Text generation is the most common OpenAI API use case: given context, the model completes or answers. Quality depends on model, sampling parameters, output constraints, and eval workflow—not prompt wording alone. This guide covers tuning shared by Chat Completions and Responses; field names follow Text generation docs.

Where it sits in the pipeline

[Business input] → [Prompt assembly + optional RAG] → [OpenAI generation] → [Parse & validate] → [Downstream]

Core parameters

ParameterRoleDev advice
modelCapability / cost / contextConfig-driven; verify docs before launch
temperatureRandomnessFactual tasks 0–0.3; creative 0.7+
top_pNucleus samplingUsually tune one of temperature or top_p
max_tokensOutput capPrevents runaway; too low truncates
stopStop sequencesUseful for templated output e.g. \n---
presence_penalty / frequency_penaltyReduce repetitionSmall tweaks on long lists

Responses API may use aliases like max_output_tokens; check API reference.

Tuning request example

curl https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "temperature": 0.2,
    "max_tokens": 600,
    "messages": [
      {"role": "system", "content": "Answer in Markdown bullets, under 200 words."},
      {"role": "user", "content": "What happens if max_tokens is too small?"}
    ]
  }'

gpt-4o-mini is an example model ID—replace with current docs.

Controlling output shape

Markdown / fixed sections

Must include:
## Conclusion (1 sentence)
## Steps (numbered, max 5)
## Notes (bullets)
No greetings or repeating the user question.

JSON / Structured Outputs

{
  "model": "gpt-4o-mini",
  "response_format": {"type": "json_object"},
  "messages": [
    {"role": "system", "content": "Output JSON only: {\"intent\":\"\",\"confidence\":0.0,\"slots\":{}}"},
    {"role": "user", "content": "I want to change my shipping address to Pudong, Shanghai"}
  ]
}

Production must-have: server-side JSON Schema validation; retry or explicit error on failure. Prefer strict schema when docs support it.

Fixed-label classification

Pick one from [SHIPPING, REFUND, OTHER]. Output the label only, no explanation.
User: {{message}}

Pair with temperature: 0 and a golden set for accuracy.

Long context and cost

StrategyNotes
Summarize historySummarize old turns with a smaller model
RAGInject relevant chunks only—see embeddings
Trim messagessystem + last N turns
Prompt caching (if offered)Repeated system prefix may reduce cost—check pricing

Input tokens include system, tool definitions, and RAG—shortening prompts often beats switching to a smaller model.

Streaming vs non-streaming

ModeBest forWatch out
stream: trueChat UIConcatenate before parsing JSON
Non-streamBatch, ETLLonger user wait

Details: API Developer Guide.

Quality evaluation workflow

  1. Golden set: 20–100 real queries + expected points (not exact strings).
  2. Automated metrics: JSON pass rate, label accuracy, keyword coverage.
  3. Human spot check: 5% of production logs weekly.
  4. Regression: rerun the set after prompt or model changes.

Versioning: Prompt Engineering.

Scenario parameter cheat sheet

Scenariotemperaturemax_tokensNotes
Support FAQ0–0.2256–512Escalate when unsure
Code explanation0.1–0.31024+Still run tests
Marketing copy0.7–0.9512–1024Human brand review
Field extraction0512JSON + schema
EN translation0.2–0.4~1.2× sourceGlossary in system

Multilingual output

  • Specify language: “Respond in English regardless of input language.”
  • Put proper nouns in system to reduce mixed-language drift.
  • Token density differs by language—measure with usage, don’t guess from character count.

Frequently asked questions

Is temperature=0 fully deterministic?

Lower randomness, but not strictly deterministic—follow official guidance.

What if max_tokens is too small?

Truncated mid-sentence or half JSON. Check finish_reason and leave headroom.

Can JSON mode still include prose?

Possible. Combine system constraints + validation; prefer strict schema.

Quality suddenly dropped?

Check: model change, RAG retrieval, prompt truncation, temperature drift.

When to consider fine-tuning?

Most teams start with prompt + RAG; evaluate Fine-tuning when you have large labeled data and fixed formats.

Official resources

Next reading

Action path

Today: Pick 5 real queries; log usage and output quality. Tomorrow: Fix temperature / max_tokens and A/B. This week: Add JSON validation + 10 golden cases in a pre-release script.

Related