Vision
Last updated:2026-08-12· 15 min read
🚀 Quick access
- ChatGPT Domestic:Open entry↗
- Mirror site:Open mirror↗
- Official ChatGPT:chatgpt.com ↗

Last updated: 2026-08-12
Overview
Vision means models process images and text together—OCR, chart interpretation, UI screenshot analysis, quality inspection, and more. OpenAI multimodal flows mainly through Responses API (preferred for new integrations) or legacy Chat Completions with image URL / base64. Vision-capable model IDs, content block field names, and vision token accounting follow Vision docs and pricing.
Task selection
| Task | Input example | Output | Preprocessing |
|---|---|---|---|
| Receipt OCR | Phone photo | JSON fields | Perspective fix, brighten |
| Document Q&A | PDF page screenshot + question | Natural language | Split by page |
| Chart summary | Dashboard screenshot | Trend bullets | Require cited numbers |
| UI review | Design mock screenshot | Diff list | Fixed viewport |
| Line QA | Product photo | pass / fail | Explicit rubric |
Skip Vision when: text-only PDF—extract text layer first; tiny dense tables—A/B dedicated OCR.
Responses API integration (recommended)
New integrations use Responses, unified with tools and conversation state:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-4o", # per vision-capable models in docs
instructions="You are a receipt OCR assistant. Use null for unreadable fields; do not guess.",
input=[{
"role": "user",
"content": [
{"type": "input_text", "text": "Extract vendor, date, total as JSON"},
{"type": "input_image", "image_url": "https://cdn.example.com/invoice.jpg"},
],
}],
)
print(response.output_text)
Field names (input_text / input_image, etc.) follow current API reference.
Legacy Chat Completions
Existing integrations with messages[].content[] and type: image_url can stay; migrate new features toward Responses.
Image input methods
| Method | Best for | Risk |
|---|---|---|
| HTTPS URL | Public CDN | User-supplied URLs need SSRF protection |
| Base64 | Intranet, direct upload | Large bodies—watch limits |
| File API + file_id | Reuse same image | See Files docs |
SSRF protection: fetch user URLs server-side only; allowlist domains; block internal and metadata IPs (169.254.x.x, etc.).
Cost control: detail and preprocessing
Vision input often converts to tile / resolution tokens. Some APIs support detail: low | high | auto:
| Level | Best for |
|---|---|
low | Scene classification, coarse description |
high | Small-text OCR, fine charts |
auto | Default strategy—see docs |
Cost checklist:
- Downscale large images to OCR-sufficient size (e.g. long edge ≤ 2048px)
- Multi-page PDF: one page per call, not 20 images at once
- Fixed-layout receipts: ROI crop first
- Cache parse results (hash → JSON) for identical images
Structured output template
[System]
Output JSON only: {"vendor":"","date":"","total":"","items":[]}
No markdown fences or explanatory prose.
[User]
(attached image)
Extract all visible fields.
Pair with Responses JSON schema / text.format or Chat response_format; validate server-side. Parameters: Text Generation guide.
Boundary with Images API
- Image Generation: text → new image
- Vision: image → understand, extract, reason
Chain: generate asset → Vision QA → Edits if fail.
Privacy and compliance
- ID cards, medical, children’s images: minimize collection; don’t log raw images
- Disclose third-party API processing; enterprise: check DPA
- Cross-border transfer and retention per org policy
Frequently asked questions
Can Vision replace dedicated OCR?
Often sufficient on clean scans; handwriting, dense tables, mixed languages need business-sample A/B.
Maximum image size?
Request body and model limits apply; compress or paginate oversized images—see docs.
Inconsistent OCR on same image?
Check temperature; OCR should use low randomness + JSON + server validation.
Scanned PDF workflow?
Rasterize to images (one per page); text PDFs prefer text extraction to save vision tokens.
Video?
Extract frames as images; streaming needs QPS and cost planning—see Realtime docs.
Official resources
Next reading
Action path
Today: Run one clear invoice through Responses API; output JSON fields. Tomorrow: Compare detail low vs high on tokens and accuracy. This week: 20-sample OCR benchmark; add SSRF-safe fetch and resize preprocessing.
Related
OpenAI Dev Overview
2026 OpenAI developer map: how ChatGPT web, Platform console, and APIs divide work—and the reading order from first call to production.
OpenAI Platform Overview
platform.openai.com console, doc navigation, Playground, usage billing, and org management—how developers find API information efficiently.
OpenAI API Quickstart
From Platform account and API key to your first OpenAI call: Responses/Completions examples, billing, rate limits, and a security checklist (2026 hands-on).
ChatGPT API Developer Guide
Production OpenAI API integration: architecture, auth, streaming, tool use, rate-limit retries, and a launch checklist.