Text Generation
Last updated:2026-08-12· 17 min read
🚀 Quick access
- ChatGPT Domestic:Open entry↗
- Mirror site:Open mirror↗
- Official ChatGPT:chatgpt.com ↗

Last updated: 2026-08-12
Overview
Text generation is the most common OpenAI API use case: given context, the model completes or answers. Quality depends on model, sampling parameters, output constraints, and eval workflow—not prompt wording alone. This guide covers tuning shared by Chat Completions and Responses; field names follow Text generation docs.
Where it sits in the pipeline
[Business input] → [Prompt assembly + optional RAG] → [OpenAI generation] → [Parse & validate] → [Downstream]
- Prompts: Prompt Engineering
- HTTP integration: API Developer Guide
- First call: API Quickstart
Core parameters
| Parameter | Role | Dev advice |
|---|---|---|
model | Capability / cost / context | Config-driven; verify docs before launch |
temperature | Randomness | Factual tasks 0–0.3; creative 0.7+ |
top_p | Nucleus sampling | Usually tune one of temperature or top_p |
max_tokens | Output cap | Prevents runaway; too low truncates |
stop | Stop sequences | Useful for templated output e.g. \n--- |
presence_penalty / frequency_penalty | Reduce repetition | Small tweaks on long lists |
Responses API may use aliases like
max_output_tokens; check API reference.
Tuning request example
curl https://api.openai.com/v1/chat/completions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"temperature": 0.2,
"max_tokens": 600,
"messages": [
{"role": "system", "content": "Answer in Markdown bullets, under 200 words."},
{"role": "user", "content": "What happens if max_tokens is too small?"}
]
}'
gpt-4o-mini is an example model ID—replace with current docs.
Controlling output shape
Markdown / fixed sections
Must include:
## Conclusion (1 sentence)
## Steps (numbered, max 5)
## Notes (bullets)
No greetings or repeating the user question.
JSON / Structured Outputs
{
"model": "gpt-4o-mini",
"response_format": {"type": "json_object"},
"messages": [
{"role": "system", "content": "Output JSON only: {\"intent\":\"\",\"confidence\":0.0,\"slots\":{}}"},
{"role": "user", "content": "I want to change my shipping address to Pudong, Shanghai"}
]
}
Production must-have: server-side JSON Schema validation; retry or explicit error on failure. Prefer strict schema when docs support it.
Fixed-label classification
Pick one from [SHIPPING, REFUND, OTHER]. Output the label only, no explanation.
User: {{message}}
Pair with temperature: 0 and a golden set for accuracy.
Long context and cost
| Strategy | Notes |
|---|---|
| Summarize history | Summarize old turns with a smaller model |
| RAG | Inject relevant chunks only—see embeddings |
| Trim messages | system + last N turns |
| Prompt caching (if offered) | Repeated system prefix may reduce cost—check pricing |
Input tokens include system, tool definitions, and RAG—shortening prompts often beats switching to a smaller model.
Streaming vs non-streaming
| Mode | Best for | Watch out |
|---|---|---|
stream: true | Chat UI | Concatenate before parsing JSON |
| Non-stream | Batch, ETL | Longer user wait |
Details: API Developer Guide.
Quality evaluation workflow
- Golden set: 20–100 real queries + expected points (not exact strings).
- Automated metrics: JSON pass rate, label accuracy, keyword coverage.
- Human spot check: 5% of production logs weekly.
- Regression: rerun the set after prompt or model changes.
Versioning: Prompt Engineering.
Scenario parameter cheat sheet
| Scenario | temperature | max_tokens | Notes |
|---|---|---|---|
| Support FAQ | 0–0.2 | 256–512 | Escalate when unsure |
| Code explanation | 0.1–0.3 | 1024+ | Still run tests |
| Marketing copy | 0.7–0.9 | 512–1024 | Human brand review |
| Field extraction | 0 | 512 | JSON + schema |
| EN translation | 0.2–0.4 | ~1.2× source | Glossary in system |
Multilingual output
- Specify language: “Respond in English regardless of input language.”
- Put proper nouns in system to reduce mixed-language drift.
- Token density differs by language—measure with
usage, don’t guess from character count.
Frequently asked questions
Is temperature=0 fully deterministic?
Lower randomness, but not strictly deterministic—follow official guidance.
What if max_tokens is too small?
Truncated mid-sentence or half JSON. Check finish_reason and leave headroom.
Can JSON mode still include prose?
Possible. Combine system constraints + validation; prefer strict schema.
Quality suddenly dropped?
Check: model change, RAG retrieval, prompt truncation, temperature drift.
When to consider fine-tuning?
Most teams start with prompt + RAG; evaluate Fine-tuning when you have large labeled data and fixed formats.
Official resources
- Text generation docs
- Structured outputs (search latest section)
- API pricing
- OpenAI Platform
Next reading
Action path
Today: Pick 5 real queries; log usage and output quality. Tomorrow: Fix temperature / max_tokens and A/B. This week: Add JSON validation + 10 golden cases in a pre-release script.
Related
OpenAI Dev Overview
2026 OpenAI developer map: how ChatGPT web, Platform console, and APIs divide work—and the reading order from first call to production.
OpenAI Platform Overview
platform.openai.com console, doc navigation, Playground, usage billing, and org management—how developers find API information efficiently.
OpenAI API Quickstart
From Platform account and API key to your first OpenAI call: Responses/Completions examples, billing, rate limits, and a security checklist (2026 hands-on).
ChatGPT API Developer Guide
Production OpenAI API integration: architecture, auth, streaming, tool use, rate-limit retries, and a launch checklist.