Qwen API & Model Studio
Last updated:2026-09-17· 17 min read
🚀 Quick access
- Qwen Max:Open entry↗
- Multi-model chat studio:Open mirror↗
- Official Qwen:chat.qwen.ai ↗

Updated: 2026-09-17. Endpoints, model IDs, prices, and rate limits follow Model Studio docs and the Bailian console on the day you ship. Placeholders below avoid short-lived model names and unit prices.
Overview
Qwen API, Alibaba Model Studio, and Bailian Qwen workflows mean wiring Qwen models into sites, scripts, bots, or internal tools—through Alibaba Cloud Model Studio (百炼 / Bailian), not the Tongyi Qwen chat app account. Production-ready integration is more than “HTTP 200 once.” You also need: keys never in the browser, timeouts and retries, schema validation, usage budgets, and redacted logs. Web chat quotas, model lists, and bills do not automatically match the API.
What this guide solves
- Separate Bailian API from Tongyi Qwen chat product accounts and billing
- Prepare Alibaba Cloud account, Bailian activation, and API Key
- Send a first request via environment variables (including OpenAI-compatible patterns)
- Design timeouts, retries, error handling, and cost controls
- Pass a pre-launch security checklist
Bailian vs chat product: know the boundary
| Dimension | Tongyi Qwen chat | Alibaba Model Studio API |
|---|---|---|
| Typical entry | Tongyi app, web chat | Bailian console |
| Primary use | Human trials, casual chat | Server-side integration, automation, batch jobs |
| Auth | Login / product account | API Key (DashScope, etc.—see docs) |
| Billing | Product plan or free tier | Pay-as-you-go / resource packs, separate bill |
| Model list | Whatever the chat UI offers | Console + API doc list |
For casual trials use Qwen Max chat or the multi-model chat studio. Product integration must call the API from a server with separate Bailian activation and billing.
Prep: account, key, billing
- Sign in to Alibaba Cloud and open the Bailian console.
- Complete Model Studio activation and billing / resource packs; set budget or usage alerts if offered.
- Create a Key under “API-KEY management” or the path shown in docs; copy immediately into a password manager or secret store.
- Tag keys by purpose (
dev-chat/prod-summarizer) so rotation after leaks is precise. - Confirm current base URL, auth header, and model ID list in Model Studio docs—do not copy expired blog snippets.
Having a Tongyi app account does not mean your API Key is ready; permissions and billing are independent.
Environment variables: secrets stay server-side
# Linux / macOS (local dev)
export DASHSCOPE_API_KEY="your-secret-here"
# Windows PowerShell (current session only; production → platform Secrets)
$env:DASHSCOPE_API_KEY = "your-secret-here"
Variable names follow official docs (often DASHSCOPE_API_KEY); this is a placeholder example.
Hard rules:
- Keys never in frontend, mobile apps, public repos, screenshots, or issues
- Use
process.env/ deployment Secrets / cloud key managers - Separate keys for dev / staging / prod
- Suspected leak → revoke immediately in console and rotate
First request (placeholders—replace from docs)
Model Studio supports an OpenAI-compatible calling style (exact path and fields in the “OpenAI compatible” doc section). Replace OFFICIAL_API_ENDPOINT and MODEL_ID_FROM_DOCS with current values.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DASHSCOPE_API_KEY"],
base_url="OFFICIAL_API_ENDPOINT", # copy from official docs
)
resp = client.chat.completions.create(
model="MODEL_ID_FROM_DOCS", # e.g. qwen-max, qwen-plus—verify in console
messages=[
{"role": "user", "content": "Explain idempotency in three bullet points"}
],
timeout=60,
)
print(resp.choices[0].message.content)
Without the OpenAI SDK, POST via requests using the REST shape in docs. Auth header names (Authorization: Bearer vs X-DashScope-Api-Key) follow live docs.
After the first success, check usage fields (if any), latency, and whether you logged the Key by mistake.
Model ID and version cautions
- Console display names (Qwen Max, Plus, Turbo) may differ from API
modelstrings—recheck before each integration. - New or preview models may need separate activation; common errors include 403 or “model not found.”
- Do not hard-code beta or date-stamped temporary IDs in production; configure fallback to stable SKUs.
- Multimodal, tool calling, and JSON modes vary by model—use the capability matrix in docs.
Reusable business prompt (inside messages)
You are an internal knowledge assistant. Answer only from the provided context.
If context is insufficient, say "Not mentioned in materials" and list missing info.
Do not invent links, regulations, or data.
context:
"""
[redacted passage]
"""
Question: [user question]
Streaming notes
Interactive UIs usually need streaming so users see partial tokens early. Bailian docs cover SSE streaming; event format and client parsing follow the official section. Notes:
- Set overall and idle timeouts for streams too
- Mid-stream disconnects should fail clearly, not spin forever
- Never hold Keys in the browser for “direct streaming”; proxy through your backend
Reliability design (production must-haves)
- Timeouts: connect + read; avoid hung requests blocking threads
- Retries: limited exponential backoff for 429 / 5xx / transient network only; do not blindly retry 401/403/400
- Idempotency: business idempotency keys on writes to prevent duplicate side effects
- Validation: schema-check required JSON; safe fallback or limited “repair” retries
- Isolation: per-user QPS caps, global rate limits, circuit breakers
- Observability: log
request_id(if returned), model name, latency, tokens, error codes—not raw PII or full private conversations - Versioning: system prompts, model names, and temperature-like params in config management
Error handling table
| Situation | Typical signal | Action |
|---|---|---|
| Auth failure | 401 / invalid key | Check env vars and Key status; no retry |
| Permission / not enabled | 403 / model unauthorized | Check console activation and balance |
| Bad params | 400 / invalid model | Fix model and fields per docs |
| Rate limit | 429 | Backoff, queue, lower concurrency, request quota increase |
| Server error | 5xx | Bounded retry + alert |
| Timeout | client timeout | Shorten context, lower max tokens, asyncify |
| Bad output | broken JSON | Validation failure path + limited retry |
Cost control
- Cap input length and max output; summarize or truncate long histories
- Cache repeatable requests (mind personalization and privacy boundaries)
- Route by tier: light models for classification, Max tier for heavy reasoning (names per docs)
- Track daily tokens and spend; set budget alerts
- Batch off-peak; avoid pointless “think again” loops
Do not hard-code “price per million tokens” in tutorials—use live Bailian pricing.
Security checklist (pre-launch)
- Keys only in server env vars or secret managers
- No keys in repos, CI logs, or frontend bundles
- User input length limits and basic filtering (per compliance)
- Redacted logs with clear retention
- Regular Key rotation; offboarding includes revocation
- Separate dev / staging / prod Keys
- Browsers and apps never hold Keys directly
Access and entry points
- Try Qwen Max: Qwen Max chat
- Multi-model studio: Multi-model chat studio
- Bailian console: bailian.console.aliyun.com
- Official docs: help.aliyun.com/zh/model-studio/
- Model family: What is Qwen?
- Coding guide: Qwen coding guide
Web entry is for human trials; API is for server integration. Model lists and billing may differ.
FAQ
Can my Tongyi app account call the API directly?
Not by default. API Keys are created separately in Bailian with their own billing; chat product quotas are independent.
Can I put the API Key in the browser?
No. Anything shipped to the browser can be extracted. Your backend should call Bailian.
OpenAI-compatible mode vs native API—which to pick?
If you already wrap OpenAI SDKs, compatible mode lowers migration cost. Greenfield projects should compare feature gaps (tools, multimodal, structured output) in docs. Endpoints and model names follow live docs—do not mix stale examples.
Model name errors?
Use the current list in console and official docs: renamed, not enabled, or typo. Keep MODEL_ID_FROM_DOCS as a reminder to recheck.
How do I control cost?
Limit context and output, cache reusable results, tier routing (Turbo/Plus/Max), budget alerts, and monitor abnormal traffic (scraped endpoints).
Key leaked—what now?
Revoke in Bailian console immediately → create new Key → audit Git history, CI logs, container env, frontend artifacts → check billing for abnormal calls.
Official resources
Next steps
Action plan
Today: Create a Key in Bailian, run one placeholder request via env vars, log usage and latency.
Tomorrow: Add timeouts, 429 backoff, and schema validation; remove any plaintext Keys from code.
This week: Set budget alerts, benchmark 10 real inputs, pick default model tier; for prompts see prompt guide.
Related
Qwen Guides Overview
2026 Qwen guides overview: learning path, official vs China access, Max/Plus/Flash tiers, entry vs Bailian API, and a five-step workflow for high-quality chats.
What is Qwen? Model Family
2026 what is Qwen: Tongyi vs Qwen vs Bailian naming, Max/Plus/Flash/Coder roles, three-step model selection, entry vs API separation, and myth busting.
How to Use Qwen in China (Complete)
2026 complete guide to using Qwen in China: official Qwen Chat vs Tongyi product vs third-party entries, step-by-step access, security checklist, and troubleshooting.
Qwen Official Entry & Signup
2026 Qwen official entry guide: verify chat.qwen.ai and qianwen.aliyun.com, complete signup and login, harden security, separate chat from Bailian API billing, and fix verification failures.