Skip to content

GPT Image Getting Started: Your First Image

Last updated:2026-08-21· 15 min read

🚀 Quick access

  • GPT Image 2 Domestic:Open entry↗
  • Text-to-image studio:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

GPT Image Getting Started: Your First Image

Updated: 2026-08-21. Available models, sizes, quality, and billing on each entry follow chatgpt.com and OpenAI Platform that day.

Introduction

This is a hands-on GPT Image start: no long model reviews—just getting one deliverable frame done. What to prepare, how to write round one, how to iterate, how to use references, and what to check before handoff. Beginners should default to GPT Image 2 (gpt-image-2, verify in docs that day). If your pipeline is still pinned to 1.5, read 1.5 selection & migrate before deciding whether to upgrade.

What this guide solves

  • Account, acceptance criteria, and asset rights before you generate
  • A step-by-step first render—not a pile of style adjectives
  • Multi-round edits that change only one thing
  • How to describe reference control, then self-check before delivery

Prep: ten minutes before you start

1. Pick an entry and log identity

Choose one to start (you can try others later, but one eval binds one entry):

EntryLinkNote
China convenience (third-party)GPT Image 2Not OpenAI official; read privacy notes first
Text-to-image studioText-to-image studioUse-oriented entry; model list follows the page that day
Official ChatGPTchatgpt.comMatches official account system
Developer docs / Platformplatform.openai.comConfirm model IDs, sizes, quality

In notes: date + URL + UI model name / docs ID. Network and account detail for China: China access guide.

2. Pick a model (beginner default)

  • Default: GPT Image 2 — flagship; edits and short titles usually steadier
  • Compat / A-B: 1.5 — only when pipelines or historical batches require it; see 1.5 guide
  • Unsure of the UI label → open the picker or Platform docs for today’s ID

Model comparison: GPT Image 2 guide.

3. Write a mini acceptance brief

Fill these four lines before opening chat:

Use: ________ (e.g. 16:9 article hero)
Subject: ________ (one object/scene; state count)
Must have: ________ (margin location, light, one-line style)
Ban: ________ (text / watermark / extra limbs / unauthorized brands)

No brief means you cannot tell which step failed.

4. Assets and rights

  • Upload only references, logos, and packaging shots you have rights to use
  • Faces, employee photos, client products need confirmed authorization
  • Do not request imitation of a living artist’s unique style

First image: step-by-step

Walk through a “ceramic table lamp hero.” Swap the subject; keep the structure.

Step 1: Lock aspect in the first sentence

Canvas first, then content. Common: article hero 16:9, square 1:1, portrait 9:16. Exact selectable sizes follow UI and Platform that day—this guide does not freeze a pixel table.

Step 2: Send round-one prompt (no text)

Use: 16:9 blog hero with ~30% right margin for a title.
Subject: one matte ceramic table lamp, cream shade, centered slightly left on a light wood desk.
Lens: eye level, soft side light, clean soft background.
Style: real product photography—no illustration, no sci-fi.
Ban: any on-image text, watermarks, frames, extra cables, warped shade.

Step 3: Composition-only accept

Ask four questions:

  1. Is the subject a table lamp (not some other fixture)?
  2. Is aspect 16:9 with right-side margin?
  3. Any obvious defects (collapsed shade, floating, broken perspective)?
  4. Any banned elements (text, watermark, random logo)?

Any fail → full redraw or rewrite the prompt—do not enter fine edits yet.

Step 4: Save the usable frame

Download or pin the composition-correct result; note v1-composition-pass. All later edits start from those constraints.

Step 5: Raise size / quality only when needed

After composition passes, raise size or quality per UI options that day. Higher tiers cannot rescue wrong composition.

Iterative edits: one variable per round

Once you have a usable first frame, obey three rules:

  1. State what stays (lens, pose, palette, background)
  2. Request one change only
  3. Recheck the brief immediately; pass before the next round

Material change

Keep the same composition, lens, lamp shape, and margin.
Only change the base material from matte ceramic to walnut.
Do not add text; do not change shade color.

Light change

Keep subject, composition, and materials unchanged.
Only change light to soft morning window light with gentler shadows.
Do not change the number of background objects.

When an edit goes wrong

  • Return to the last still-correct frame and re-instruct
  • Or start a new chat: paste the brief + a short description of the good frame, attach it as reference if needed
  • Avoid stacking ten contradictory edits in a drifting long thread

On multi-round ability, 2 usually beats grinding 1.5. See GPT Image 2 guide.

How to use references

References are not “drop in and hope”—declare control scope.

Image 1: lamp silhouette and proportions (geometry must be close).
Image 2: palette only (warm wood + cream)—do not copy Image 2 clutter.
Generate: 16:9 lamp on light wood desk with right margin; no text or watermark.

Practical rules

PracticeNote
Say what each image controlsGeometry / palette / scene mood / character identity—separate lines
Limit countBeginners start with 1; strong multi-ref consistency prefers GPT Image 2
Priority on conflict“Geometry from Image 1; color from Image 2”
CopyrightUnauthorized images are not a commercial delivery basis

Products and logos: after generate, compare item-by-item to brand source files—“looks close” is not a pass.

Done checklist (delivery)

Before sending externally, tick:

  • Use and aspect correct; crop/title safe margins enough
  • Subject type and count correct; no obvious defects or extra objects
  • Iteration did not invent text, watermarks, or wrong logos
  • If references used: consistency acceptable; rights traceable
  • Required on-image text proofed character-by-character; zero-tolerance copy moved to later typeset
  • Checked at target display size and selected quality (thumbnail + full)
  • Logged entry, model name/ID (that day), prompt summary for reproduction
  • Commercial release: rights confirmed; sensitive info (faces, addresses, unpublished products) handled

Fast path

FAQ

Start on 1.5 or 2?

Learning the flow: use 2—fight fewer legacy limits. Must match historical pipelines: read 1.5 selection, then decide on dual-run A-B.

Why does a long prompt make the image messier?

Long prompts collide. Lock use, subject, aspect, and bans first; add one style layer at a time. System writing: prompt playbook.

On-image text keeps failing—what now?

Short titles: quote them and fix in a dedicated round. Legal microcopy, dense multilingual stacks, QR codes: generate blank, typeset in design tools. 2 is usually better than 1.5 for short titles—still final-check character-by-character.

Can I use official ChatGPT from China directly?

Depends on your network and account conditions; policy can change. Use labeled China convenience entries for learning first, then connect official as needed. See China access.

What does it cost? Is it free?

Follow each entry and official docs that day; this tutorial does not hard-code unit prices or quotas.

Official resources

Further reading

Summary

Beginner success is not “one perfect poster.” It is: brief → round one validates composition → one change per round → references declare control → delivery checklist. Smooth that loop before style libraries and API; model, size, quality, and price always follow official docs that day.

Related