Skip to content

Vision

Last updated:2026-08-12· 15 min read

🚀 Quick access

  • ChatGPT Domestic:Open entry↗
  • Mirror site:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

Vision

Last updated: 2026-08-12

Overview

Vision means models process images and text together—OCR, chart interpretation, UI screenshot analysis, quality inspection, and more. OpenAI multimodal flows mainly through Responses API (preferred for new integrations) or legacy Chat Completions with image URL / base64. Vision-capable model IDs, content block field names, and vision token accounting follow Vision docs and pricing.

Task selection

TaskInput exampleOutputPreprocessing
Receipt OCRPhone photoJSON fieldsPerspective fix, brighten
Document Q&APDF page screenshot + questionNatural languageSplit by page
Chart summaryDashboard screenshotTrend bulletsRequire cited numbers
UI reviewDesign mock screenshotDiff listFixed viewport
Line QAProduct photopass / failExplicit rubric

Skip Vision when: text-only PDF—extract text layer first; tiny dense tables—A/B dedicated OCR.

New integrations use Responses, unified with tools and conversation state:

from openai import OpenAI

client = OpenAI()
response = client.responses.create(
    model="gpt-4o",  # per vision-capable models in docs
    instructions="You are a receipt OCR assistant. Use null for unreadable fields; do not guess.",
    input=[{
        "role": "user",
        "content": [
            {"type": "input_text", "text": "Extract vendor, date, total as JSON"},
            {"type": "input_image", "image_url": "https://cdn.example.com/invoice.jpg"},
        ],
    }],
)
print(response.output_text)

Field names (input_text / input_image, etc.) follow current API reference.

Legacy Chat Completions

Existing integrations with messages[].content[] and type: image_url can stay; migrate new features toward Responses.

Image input methods

MethodBest forRisk
HTTPS URLPublic CDNUser-supplied URLs need SSRF protection
Base64Intranet, direct uploadLarge bodies—watch limits
File API + file_idReuse same imageSee Files docs

SSRF protection: fetch user URLs server-side only; allowlist domains; block internal and metadata IPs (169.254.x.x, etc.).

Cost control: detail and preprocessing

Vision input often converts to tile / resolution tokens. Some APIs support detail: low | high | auto:

LevelBest for
lowScene classification, coarse description
highSmall-text OCR, fine charts
autoDefault strategy—see docs

Cost checklist:

  • Downscale large images to OCR-sufficient size (e.g. long edge ≤ 2048px)
  • Multi-page PDF: one page per call, not 20 images at once
  • Fixed-layout receipts: ROI crop first
  • Cache parse results (hash → JSON) for identical images

Structured output template

[System]
Output JSON only: {"vendor":"","date":"","total":"","items":[]}
No markdown fences or explanatory prose.

[User]
(attached image)
Extract all visible fields.

Pair with Responses JSON schema / text.format or Chat response_format; validate server-side. Parameters: Text Generation guide.

Boundary with Images API

  • Image Generation: text → new image
  • Vision: image → understand, extract, reason

Chain: generate asset → Vision QA → Edits if fail.

Privacy and compliance

  • ID cards, medical, children’s images: minimize collection; don’t log raw images
  • Disclose third-party API processing; enterprise: check DPA
  • Cross-border transfer and retention per org policy

Frequently asked questions

Can Vision replace dedicated OCR?

Often sufficient on clean scans; handwriting, dense tables, mixed languages need business-sample A/B.

Maximum image size?

Request body and model limits apply; compress or paginate oversized images—see docs.

Inconsistent OCR on same image?

Check temperature; OCR should use low randomness + JSON + server validation.

Scanned PDF workflow?

Rasterize to images (one per page); text PDFs prefer text extraction to save vision tokens.

Video?

Extract frames as images; streaming needs QPS and cost planning—see Realtime docs.

Official resources

Next reading

Action path

Today: Run one clear invoice through Responses API; output JSON fields. Tomorrow: Compare detail low vs high on tokens and accuracy. This week: 20-sample OCR benchmark; add SSRF-safe fetch and resize preprocessing.

Related