Skip to content
Main Site News Console

Billing and Quotas

Billing Method

  • Pay-as-you-go: conversations/embeddings are billed by token usage (input and output are priced separately), images are billed by count and specification, and videos are billed by duration and specification;
  • No plans, no monthly fees, no lock-in: use as much as you top up, and the balance remains valid long term;
  • See the console Pricing page for each model’s unit price; when prices change due to upstream adjustments, they will be announced in the Changelog.

Quotas and Tokens

  • Account balance is a shared pool, and each token can be configured with its own quota limit; once exceeded, requests made with that token are rejected without affecting other tokens;
  • Tokens can also be set with expiration times and model groups, making it easy to isolate by project/environment.

Usage Lookup

The console Logs page provides details for each call: model, input/output tokens, billed quota, and latency; the Dashboard can summarize by day/by model and supports export for reconciliation.

FAQ

QuestionDescription
Is streaming interruption still billed?The portion already generated is billed based on actual usage
Is a failed request charged?Requests that return errors from the gateway or upstream are not charged; failed generation tasks will վերադարձ the prepaid quota
How are tokens estimated?In Chinese, roughly 1 character ≈ 1–2 tokens; you can verify precisely with the usage field in the response
How do I top up?In the console Wallet page, online payment is supported; for enterprise wire transfer, please contact customer support