Providers

Nine labs behind one key.

Every upstream the gateway routes to, what it is good for, and which currency it prices in. Each card links to that provider's own documentation so you can verify anything here at the source.

CNY Docs ↗

Moonshot AI

Long-context frontier models with native vision and configurable reasoning depth. Kimi K3 offers up to 1M context and is billed in CNY, with a substantially lower cache-hit input rate.

Kimi K3
CNY Docs ↗

Zhipu AI

GLM-5.2 pairs 1M context with up to 128K output, function calling and structured output — the combination most development workflows actually need.

GLM-5.2
CNY Docs ↗

Alibaba

Text and video in one house. Qwen3.8 Max Preview is a 1M-context reasoning model still in preview; Wan 2.7 generates video billed per generated second at 720P and 1080P.

Qwen3.8 Max Preview · Wan 2.7 T2V
USD Docs ↗

DeepSeek

1M context, up to 384K output, tool calls and thinking modes — priced in USD and with one of the steepest cache-hit discounts of any provider on this page.

DeepSeek V4 Pro
USD Docs ↗

BytePlus

A fast text model and an image model in the same family. Seedream 5.0 Pro gives the first image free, then bills per image with a higher rate only above 2.36M pixels.

Seed 2.1 Turbo · Seedream 5.0 Pro
CNY Docs ↗

HappyHouse

Text-to-video generation billed per generated second, at 720P and 1080P — an alternative to Wan when you want a second creative pipeline on the same key.

HappyHorse 1.1 T2V
USD Docs ↗

OpenAI

The widest surface on the platform: 40+ models spanning flagship chat, reasoning, image, transcription and realtime audio, all under the schema the rest of the gateway is modelled on.

GPT-5.6 Sol / Terra / Luna · o-series · GPT-Image · Realtime
USD Docs ↗

Anthropic

Claude Sonnet 5 currently carries an introductory input rate through 31 Aug 2026. Six earlier models remain routable for teams pinned to a specific version.

Claude Sonnet 5 · Haiku 4.5 · Opus 4.6
USD Docs ↗

Google

Ten models including image generation, a Computer Use preview and the Live API for streaming audio and video, with long-context tiers above 200K input tokens.

Gemini 3.5 Flash · 3.1 Pro · Computer Use · Live API
Why it matters

Nine providers is not a
vanity number.

Somewhere to fall back to

A fallback list is only useful if the alternatives are genuinely comparable. Breadth across providers is what makes failover a real answer rather than a slower error.

Price pressure

When a task runs acceptably on three different models, you get to pick on cost. That only works if all three are already one config change away.

No single point of policy

Providers change terms, regions and availability independently. Spanning several of them means one provider's decision is not automatically your outage.