Skip to content

Prompt Engineering

Last updated:2026-08-12· 18 min read

🚀 Quick access

  • ChatGPT Domestic:Open entry↗
  • Mirror site:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

Prompt Engineering

Last updated: 2026-08-12

Overview

In production, prompts are versioned code: they define output format, safety boundaries, and failure modes. This guide covers engineering prompts—Git-managed, regression-tested, observable—not magic phrases. Use with Text Generation parameters and API integration.

Four-layer prompt structure

[Static system]     Role, policy, schema (low change frequency)
[Dynamic context]   RAG, user profile (per request)
[Few-shot]          Optional format examples
[user]              End-user input
LayerChange frequencyStorage
system templateLowGit / CMS + version id
RAG contextHighVector store + citation ids
userPer sessionDatabase
eval setMediumJSONL in repo

System message design

  1. Clear boundaries: what you can do, cannot do, and how to say “unsure.”
  2. Parseable format: fixed JSON keys, enums, table column names.
  3. Refusal policy: don’t invent outside sources; escalate sensitive cases.
  4. Align with tools: function names and params match system descriptions.

RAG assistant system template

You are an internal knowledge assistant.
Rules:
1) Answer only from "context"; do not use facts outside context.
2) If context is insufficient, reply "Not mentioned in the materials" and list what's missing.
3) Output Markdown; cite with [paragraph index].
4) Never output keys, internal URLs, or unreleased financial data.

Few-shot: when it helps

Good fitPoor fit
Fixed format (classification, extraction)Open-domain factual Q&A
Brand tone (1–2 examples)Too many examples filling context
{
  "messages": [
    {"role": "system", "content": "Classify tickets as BUG|FEATURE|QUESTION. Output label only."},
    {"role": "user", "content": "Payment page keeps spinning"},
    {"role": "assistant", "content": "BUG"},
    {"role": "user", "content": "Can you add dark mode"},
    {"role": "assistant", "content": "FEATURE"},
    {"role": "user", "content": "{{actual_input}}"}
  ]
}

Examples should match real distribution; idealized samples cause overfitting.

Chain-of-thought in production

Complex tasks may ask for “analyze then conclude,” but:

  • Hide reasoning from users: return final JSON only; log intermediate steps if compliant.
  • Cost: long reasoning increases output tokens.
  • Eval: compare direct JSON vs stepwise on golden set.
List constraints internally in <analysis> (do not return to frontend);
output only the final JSON object.

Gateway can strip <analysis>; or use documented reasoning models if applicable.

Tool-call prompt alignment

Stable tool use depends on:

  • description stating when to call
  • Minimal, explicit parameter schema
  • system forbidding “pretend you already called”
{
  "type": "function",
  "function": {
    "name": "lookup_order",
    "description": "Use when user gives an order number and asks for status",
    "parameters": {
      "type": "object",
      "properties": {
        "order_id": {"type": "string", "description": "Numeric order id only"}
      },
      "required": ["order_id"]
    }
  }
}

Agent scenarios: Agents SDK, Responses API.

Version management and release

prompt_id: order_intent_v2
model: gpt-4o-mini
temperature: 0
changelog: "Separate change-address vs change-phone intents"
golden_set: tests/prompts/order_intent.jsonl
  1. Change prompt → bump version
  2. Run golden set (automated + spot check)
  3. Canary 5% → full rollout

Evaluation metrics

TaskMetrics
ClassificationAccuracy, confusion matrix
JSON extractionSchema pass rate, field F1
SummarizationKey-point coverage (rubric)
SupportHallucination rate, escalation rate

Use business rubrics, not BLEU alone.

Defending against prompt injection

User text may conflict with system. Mitigations:

  • System declares ignoring “leak system / change rules” instructions
  • Input length limits and pattern detection
  • Server-side permission checks before tool execution
  • Never put secrets or connection strings in prompts

ChatGPT web vs API

WebAPI
Partial system hiddenFull control of messages
Model policy follows product updatesLock model + parameters
Hard to batch regressionCI runs golden set

Web is for exploration; finalize in API with eval. Web behavior may differ from API.

Frequently asked questions

Can I put rules in user instead of system?

Some APIs allow it; still prefer system for gateway management and caching.

Are longer prompts always better?

No. Bloat wastes tokens, dilutes focus, and increases conflicting instructions.

Must I retest when changing model?

Yes. Format adherence, tools, and multilingual behavior vary by model.

Are Assistants instructions the same as system?

Conceptually similar, but with files and tools—and possible migration; see Assistants guide.

How to manage multilingual prompts?

Maintain system per locale, or one system specifying output language + glossary.

Official resources

Next reading

Action path

Today: Export production system to Git with a prompt_id. Tomorrow: Write 10 golden cases + an automated validation script. This week: Require golden regression before prompt changes; log model / temperature / prompt_id.

Related