Gemini API Guide
Last updated:2026-08-12· 16 min read
🚀 Quick access
- ChatGPT Domestic:Open entry↗
- Mirror site:Open mirror↗
- Official ChatGPT:chatgpt.com ↗

Last updated: 2026-09-17. Flagship Flash: Gemini 3.8 Flash (gemini-3.8-flash).
Introduction
To embed Gemini in products, scripts, or automation, the API is the path that supports version lock and metered billing. This guide starts with Google AI Studio key setup, walks through your first HTTP call, and covers model choice, key security, and launch checks. Model IDs, quotas, pricing, and regional availability follow Google AI documentation—this guide does not list fixed prices.
API vs web: why use the API?
| Item | Web Gemini | Gemini API |
|---|---|---|
| Integration | Manual copy-paste | Programmatic calls |
| Model lock | Product auto-routing | Request specifies model |
| Billing | Subscription (per account) | Per token/request (per official site) |
| Logs and monitoring | Browser history | Your logs and tracing |
| Best for | Personal productivity | Product features, batch jobs, agents |
API key setup: step by step
1) Open Google AI Studio
- Visit aistudio.google.com.
- Sign in with Google; accept developer terms on first use.
- If prompted to enable Cloud billing, follow official guidance (free tier and paid thresholds per site).
2) Create an API key
- Go to Get API key or API Keys in the sidebar.
- Create a new key linked to a Google Cloud project (create one via the wizard if needed).
- Copy and store immediately—keys may show only once or on a new page.
3) Secure storage
# Example: environment variable (never commit to Git)
export GEMINI_API_KEY="your_key_here"
- Never put keys in frontend JavaScript or public repos.
- Use Secret Manager or CI secrets in production.
- Rotate keys regularly; revoke immediately if leaked.
First call: minimal REST example
This uses the generateContent endpoint; replace MODEL_ID with a current ID from model list.
curl "https://generativelanguage.googleapis.com/v1beta/models/MODEL_ID:generateContent?key=${GEMINI_API_KEY}" \
-H 'Content-Type: application/json' \
-d '{
"contents": [{
"parts": [{"text": "Explain Gemini API in three sentences, Simplified Chinese."}]
}]
}'
On success, JSON returns with text in candidates[0].content.parts[0].text.
Python SDK example (recommended)
import os
from google import genai
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
response = client.models.generate_content(
model="MODEL_ID", # per official model list
contents="List three API call cautions in a Markdown table, Chinese.",
)
print(response.text)
Install and latest SDK usage: official quickstart.
Model selection quick reference
| Scenario | Suggested tier | Notes |
|---|---|---|
| Complex reasoning, long-chain analysis | Ultra family | Higher cost—per official pricing |
| Daily chat, general generation | Pro family | Default for many products |
| High QPS, low latency, large batches | Flash family | Preprocessing and simple classification |
| Multimodal (image + text) | Vision-capable models | Check input modalities on model card |
Production tip: Store GEMINI_MODEL in env vars for no-downtime tier switches; monitor token usage and error rates.
Common API capabilities
| Capability | Use | Doc direction |
|---|---|---|
generateContent | Single/multi-turn text and multimodal | Core chat completion |
| Structured output / JSON mode | Reliable field parsing | Official structured output guide |
| Function calling | Agents, tool chains | Schema for tool calls |
| Files API | Large files, PDF | Upload then reference file URI |
| Embeddings | Semantic search, RAG | Separate embedding models |
Parameter names and limits change with API versions—integrate against current ai.google.dev docs.
Production checklist
- Keys: Server-side only; IP limits or project quotas if supported.
- Timeout and retry: Exponential backoff on 429/5xx; set reasonable
timeout. - Input length: Summarize very long context before generation to control cost.
- Output review: Keyword filters and spot checks for sensitive flows.
- Logs: Record
model, token usage, latency—do not log raw user PII. - Deprecation watch: Follow Google model deprecation notices; plan migration windows.
3 copy-ready developer system prompts
Use in API system or first user message:
1) JSON structured output
Output valid JSON only—no Markdown code fences.
schema: {"title": string, "bullets": string[], "confidence": "high"|"medium"|"low"}
Fill from user input; use low confidence when uncertain.
2) RAG Q&A (with context)
Answer only from the "reference material" below; for anything not mentioned, reply "not stated in material."
Material:
---
{context}
---
Question: {question}
3) Code review
Review the code below; output: ① critical issues (if any) ② suggested improvements ③ local test commands to run.
Do not invent APIs; reply in Chinese.
Code:
{code}
Notes for developers in China
- API requests originate from server region—must comply with Google Cloud / AI supported regions.
- If local access is unstable, deploy the call layer on overseas VPS or Cloud Run.
- Local debugging may need compliant network access—never commit API keys to public GitHub.
- Web experience: Official Entry (China); ChatGPT Platform or ChatGPT domestic access for product UX reference only—not API equivalents.
Frequently asked questions
Is AI Studio key the same as Vertex AI?
Not exactly. Consumer Google AI Studio / Gemini API vs Vertex AI (enterprise GCP) differ in auth, billing, and SLA. Prototypes favor AI Studio; large GCP customers may evaluate Vertex—see official comparison.
How much free quota is there?
Free tier limits change—see pricing page. Beyond free tier, pay-as-you-go or billing account required.
How to handle 429 RESOURCE_EXHAUSTED?
Quota or rate limit hit. Lower concurrency, try Flash, request quota increase, or add caching; check for exposed keys causing abuse.
Can I use Gemini API in commercial products?
Read Google API terms and generative AI use policy; some industries have extra limits—legal review before launch.
Official resources
Next reading
- What is Google Gemini?
- Gemini Prompt Engineering
- Gemini 3.8 Flash: Agent & Coding Shift vs 3.5
- Gemini 3.5 Flash Review
Action path
Today: Create a key in AI Studio; run first generateContent via curl or Python. Tomorrow: Move the key to env vars; remove plaintext from code. This week: Wire one developer system prompt into a prototype and read 3.8 Flash plus 3.5 Flash Review for cost-tier evaluation.
Related
Gemini Overview
2026 Gemini starter map: product matrix, web vs API roles, multimodal limits, and steps for your first high-quality conversation with copy-ready prompts.
What is Google Gemini?
2026 deep dive into the Gemini model family: how Ultra, Pro, and Flash are positioned, how to choose, and how to compare with GPT and Claude.
Gemini Signup & Usage
2026 step-by-step: Google account setup, Gemini Chinese conversation settings, common features, and a beginner practice checklist.
Gemini Official Entry (China)
2026 authoritative guide: gemini.google.com official entry, domain verification, network context in China, and safe alternative paths.