Moonshot AI
Long-context frontier models with native vision and configurable reasoning depth. Kimi K3 offers up to 1M context and is billed in CNY, with a substantially lower cache-hit input rate.
Every upstream the gateway routes to, what it is good for, and which currency it prices in. Each card links to that provider's own documentation so you can verify anything here at the source.
Long-context frontier models with native vision and configurable reasoning depth. Kimi K3 offers up to 1M context and is billed in CNY, with a substantially lower cache-hit input rate.
GLM-5.2 pairs 1M context with up to 128K output, function calling and structured output — the combination most development workflows actually need.
Text and video in one house. Qwen3.8 Max Preview is a 1M-context reasoning model still in preview; Wan 2.7 generates video billed per generated second at 720P and 1080P.
1M context, up to 384K output, tool calls and thinking modes — priced in USD and with one of the steepest cache-hit discounts of any provider on this page.
A fast text model and an image model in the same family. Seedream 5.0 Pro gives the first image free, then bills per image with a higher rate only above 2.36M pixels.
Text-to-video generation billed per generated second, at 720P and 1080P — an alternative to Wan when you want a second creative pipeline on the same key.
The widest surface on the platform: 40+ models spanning flagship chat, reasoning, image, transcription and realtime audio, all under the schema the rest of the gateway is modelled on.
Claude Sonnet 5 currently carries an introductory input rate through 31 Aug 2026. Six earlier models remain routable for teams pinned to a specific version.
Ten models including image generation, a Computer Use preview and the Live API for streaming audio and video, with long-context tiers above 200K input tokens.
A fallback list is only useful if the alternatives are genuinely comparable. Breadth across providers is what makes failover a real answer rather than a slower error.
When a task runs acceptably on three different models, you get to pick on cost. That only works if all three are already one config change away.
Providers change terms, regions and availability independently. Spanning several of them means one provider's decision is not automatically your outage.
Message any of the contacts below.
Scan any QR code to add a sales contact.