Skip to content

Gemini API Guide

Last updated:2026-08-12· 16 min read

🚀 Quick access

  • ChatGPT Domestic:Open entry↗
  • Mirror site:Open mirror↗
  • Official ChatGPT:chatgpt.com ↗

Gemini API Guide

Last updated: 2026-09-17. Flagship Flash: Gemini 3.8 Flash (gemini-3.8-flash).

Introduction

To embed Gemini in products, scripts, or automation, the API is the path that supports version lock and metered billing. This guide starts with Google AI Studio key setup, walks through your first HTTP call, and covers model choice, key security, and launch checks. Model IDs, quotas, pricing, and regional availability follow Google AI documentation—this guide does not list fixed prices.

API vs web: why use the API?

ItemWeb GeminiGemini API
IntegrationManual copy-pasteProgrammatic calls
Model lockProduct auto-routingRequest specifies model
BillingSubscription (per account)Per token/request (per official site)
Logs and monitoringBrowser historyYour logs and tracing
Best forPersonal productivityProduct features, batch jobs, agents

API key setup: step by step

1) Open Google AI Studio

  1. Visit aistudio.google.com.
  2. Sign in with Google; accept developer terms on first use.
  3. If prompted to enable Cloud billing, follow official guidance (free tier and paid thresholds per site).

2) Create an API key

  1. Go to Get API key or API Keys in the sidebar.
  2. Create a new key linked to a Google Cloud project (create one via the wizard if needed).
  3. Copy and store immediately—keys may show only once or on a new page.

3) Secure storage

# Example: environment variable (never commit to Git)
export GEMINI_API_KEY="your_key_here"
  • Never put keys in frontend JavaScript or public repos.
  • Use Secret Manager or CI secrets in production.
  • Rotate keys regularly; revoke immediately if leaked.

First call: minimal REST example

This uses the generateContent endpoint; replace MODEL_ID with a current ID from model list.

curl "https://generativelanguage.googleapis.com/v1beta/models/MODEL_ID:generateContent?key=${GEMINI_API_KEY}" \
  -H 'Content-Type: application/json' \
  -d '{
    "contents": [{
      "parts": [{"text": "Explain Gemini API in three sentences, Simplified Chinese."}]
    }]
  }'

On success, JSON returns with text in candidates[0].content.parts[0].text.

import os
from google import genai

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

response = client.models.generate_content(
    model="MODEL_ID",  # per official model list
    contents="List three API call cautions in a Markdown table, Chinese.",
)
print(response.text)

Install and latest SDK usage: official quickstart.

Model selection quick reference

ScenarioSuggested tierNotes
Complex reasoning, long-chain analysisUltra familyHigher cost—per official pricing
Daily chat, general generationPro familyDefault for many products
High QPS, low latency, large batchesFlash familyPreprocessing and simple classification
Multimodal (image + text)Vision-capable modelsCheck input modalities on model card

Production tip: Store GEMINI_MODEL in env vars for no-downtime tier switches; monitor token usage and error rates.

Common API capabilities

CapabilityUseDoc direction
generateContentSingle/multi-turn text and multimodalCore chat completion
Structured output / JSON modeReliable field parsingOfficial structured output guide
Function callingAgents, tool chainsSchema for tool calls
Files APILarge files, PDFUpload then reference file URI
EmbeddingsSemantic search, RAGSeparate embedding models

Parameter names and limits change with API versions—integrate against current ai.google.dev docs.

Production checklist

  1. Keys: Server-side only; IP limits or project quotas if supported.
  2. Timeout and retry: Exponential backoff on 429/5xx; set reasonable timeout.
  3. Input length: Summarize very long context before generation to control cost.
  4. Output review: Keyword filters and spot checks for sensitive flows.
  5. Logs: Record model, token usage, latency—do not log raw user PII.
  6. Deprecation watch: Follow Google model deprecation notices; plan migration windows.

3 copy-ready developer system prompts

Use in API system or first user message:

1) JSON structured output

Output valid JSON only—no Markdown code fences.
schema: {"title": string, "bullets": string[], "confidence": "high"|"medium"|"low"}
Fill from user input; use low confidence when uncertain.

2) RAG Q&A (with context)

Answer only from the "reference material" below; for anything not mentioned, reply "not stated in material."
Material:
---
{context}
---
Question: {question}

3) Code review

Review the code below; output: ① critical issues (if any) ② suggested improvements ③ local test commands to run.
Do not invent APIs; reply in Chinese.
Code:
{code}

Notes for developers in China

  • API requests originate from server region—must comply with Google Cloud / AI supported regions.
  • If local access is unstable, deploy the call layer on overseas VPS or Cloud Run.
  • Local debugging may need compliant network access—never commit API keys to public GitHub.
  • Web experience: Official Entry (China); ChatGPT Platform or ChatGPT domestic access for product UX reference only—not API equivalents.

Frequently asked questions

Is AI Studio key the same as Vertex AI?

Not exactly. Consumer Google AI Studio / Gemini API vs Vertex AI (enterprise GCP) differ in auth, billing, and SLA. Prototypes favor AI Studio; large GCP customers may evaluate Vertex—see official comparison.

How much free quota is there?

Free tier limits change—see pricing page. Beyond free tier, pay-as-you-go or billing account required.

How to handle 429 RESOURCE_EXHAUSTED?

Quota or rate limit hit. Lower concurrency, try Flash, request quota increase, or add caching; check for exposed keys causing abuse.

Can I use Gemini API in commercial products?

Read Google API terms and generative AI use policy; some industries have extra limits—legal review before launch.

Official resources

Next reading

Action path

Today: Create a key in AI Studio; run first generateContent via curl or Python. Tomorrow: Move the key to env vars; remove plaintext from code. This week: Wire one developer system prompt into a prototype and read 3.8 Flash plus 3.5 Flash Review for cost-tier evaluation.

Related