Skip to content

GPT Image Prompt Playbook

Last updated:2026-08-21· 16 min read

🚀 Quick access

  • GPT Image 2 Domestic:Open entry↗
  • Text-to-image studio:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

GPT Image Prompt Playbook

Updated: 2026-08-21. Model capability, UI options, and available models follow ChatGPT and OpenAI Images docs that day.

Introduction

Searches for “GPT Image prompt” or “gpt-image-2 prompt” usually fail on adjective piles. Delivery quality jumps when the prompt is an acceptable art brief: use, canvas, subject, composition, light, materials, on-image text, and bans. This guide gives a fixed structure, three paste-ready templates, and stable patterns for edits and text—so you stop shipping “looks fine but cannot go live.” Model choice: GPT Image 2 complete guide. First ops: Getting started.

What this guide solves

  • Turn empty prompts into reusable seven-part image specs
  • Get paste-ready templates for product stills, poster/social, and local edits
  • Raise short on-image copy hit rate—and know when to typeset later
  • Use negatives and single-variable iteration to control style drift and subject warp

Write prompts as art briefs

Suggested order (you can compress; do not scramble the logic):

  1. Use and canvas: hero / PDP / poster; aspect and rough resolution tier
  2. Subject: what, count, pose, must-keep brand or product traits
  3. Composition and lens: center / thirds, margins, viewpoint, focal feel
  4. Light and palette: key light direction, contrast, brand or banned colors
  5. Material and style: photo / 3D / illustration / flat infographic
  6. On-image text: quote short lines; state language and placement
  7. Bans: watermarks, extra limbs, cluttered backgrounds, unauthorized logos

Empty: “draw a nice coffee shop” → hard to accept.
Spec: “16:9 blog hero; latte with latte art slightly left; shallow 50mm; window light; warm cream and deep green; photo-real product look; no text; no logos or watermarks” → iterable.

In ChatGPT, write the seven parts as one coherent instruction; on Platform, version the same block as prompt. Family differences: family explainer.

Three paste-ready templates

Product still life

Use: ecommerce PDP main, 1:1.
Subject: [full product name and material], fully visible, slight 3/4 angle.
Background: clean [color] seamless, soft contact shadow on the ground.
Light: soft top-side key; highlights not blown.
Style: commercial product photography with real material texture.
Text: none.
Ban: hands, damaged packaging, watermarks, extra props stealing focus.

Poster / social cover

Use: social cover, 16:9.
Theme: [one-sentence theme].
Composition: large left margin for title; subject on the right.
Subject: [person or object], mood [keywords].
Style: [flat illustration / editorial photo / 3D render].
On-image text (must be exact): "[short title]", English, bold, upper-left safe zone.
Palette: [primary] + [secondary].
Ban: legal microcopy, QR codes, extra logos.

Local edit (multi-round)

Keep the previous lens, pose, palette, and background unchanged.
Change only: [single variable, e.g. swap the cup for a closed notebook].
Do not redraw the whole frame; do not change subject identity or light direction.

Change one variable per round—or you cannot tell which sentence broke the image. When uploading references, state whether they control packaging geometry, palette, or character identity so the model does not freestyle. Style capture: styles & scenes.

On-image text that stays readable

  • Keep copy short; quote exact strings.
  • State language, case, rough hierarchy, and placement (upper left / bottom safe zone).
  • Zero-tolerance spelling or complex trademark glyphs: generate blank, typeset in design tools.
  • Legal fine print, dense tables, QR codes: do not expect one perfect generate—overlay later.
  • Mixed languages: quote each string separately and set primary/secondary hierarchy to cut garble and misplacement.

GPT Image usually handles short titles better than long paragraphs. If text is a hard acceptance metric, split “plate generate” and “typeset overlay”—faster than one-shot perfection.

Negatives and diagnosis table

SymptomLikely causeNext round
Style jumpsToo many conflicting adjectivesCut to 1 main style + 1 material reference
Subject warpsWeak constraints or too many changesLock lens/pose; change one object
Soft textToo many characters or too smallShorten copy or blank + typeset
Brand mismatchNo reference or missing must-have traitsUpload ref; say “keep packaging geometry and color blocks”
Dirty backgroundMissing bansState “clean background, no clutter, no watermark”
Weird limbsComplex pose without identity lockSimplify pose, or change motion in steps

Negatives are not “longer is better.” Write a checkable list (no watermark, no extra fingers, no warped logos)—more useful than empty “low quality, blurry, bad anatomy” piles. If the UI has a separate Avoid field, put bans there; otherwise end the prompt with them.

Done checklist

  • Prompt includes use, aspect, subject, light, style, bans
  • On-image text proofed character-by-character (or later typeset decided)
  • This round changed one variable only
  • Reference rights confirmed
  • Limbs and artifacts checked at thumbnail and target size
  • Winning version saved with prompt and model tier

Fast path / access

Third-party accounts, logs, and routing may differ; do not upload sensitive assets. More paths: China access.

FAQ

Are longer prompts always better?

No. Long colliding adjectives lower control. Hard constraints first, style second—beats twenty synonyms.

Should I write prompts in English?

English works well for photo and material terms; you can mix languages on key nouns. Compare the same brief on your current entry and keep what wins.

Do negatives need a separate field?

Some UIs have Avoid; otherwise end the same prompt with bans. What matters is checkable: no watermark, no extra fingers, no warped logos.

Same prompt, different every time—normal?

Yes. For reproducibility, lock model tier, aspect, references, and core sentences, and save winning version IDs. API cases: API & Platform.

Do GPT Image 1.5 and 2 use the same prompt shape?

Structure can be shared; fine texture, complex text, and edit stability may differ by model. Migration notes: GPT Image 1.5 selection.

Official resources

Further reading

Summary

GPT Image prompting is specification: use and canvas first, subject and composition acceptable, text short and exact, bans explicit, one variable per iteration. Run templates first, then crystallize team style cards; move to API and cost observability when you need batch and product wiring.

Related