Pricing

Provider reference prices,
published in full.

No markup table hidden behind a sales call. Every rate below is the provider's own reference price, with a link to the source and the date we last checked it.

No subscription

No seat licence, no monthly platform fee, no minimum commitment. Top up a balance and draw against it.

One invoice

Chinese and Western models, text and media, all settle against the same balance in one usage ledger.

Alipay / WeChat Pay

No overseas card or bank account required. Talk to sales if your team needs invoicing.

These are provider reference prices, last checked 22 Jul 2026. Promotions, regions and model availability change — the linked official pages are always authoritative. Confirm your effective rate in the console before planning a budget around it.
Billing

How a call turns into a charge.

What exactly is a token?+
A token is roughly three to four characters of English, or often a single Chinese character. Both your prompt and the model's reply are counted, at separate input and output rates.
What is a cache hit rate?+
Several providers charge a much lower input rate for prompt prefixes they have already processed. If your agent resends the same system prompt and tool definitions every turn, most of that input bills at the cache-hit rate rather than the full rate.
Why are some models priced in CNY and others in USD?+
Because the upstream provider prices them that way. We publish each rate in its native currency rather than converting it, so you can check it against the provider's own page without arguing about an exchange rate.
What does “long context” pricing mean?+
Some models step up to a higher rate once a request exceeds a threshold — commonly 200K tokens of input. Rows where this applies state the threshold and the higher rate.

Need a volume or invoicing arrangement?

If you are moving sustained production traffic, talk to sales before you top up.