Billing and Quotas
Billing Method
- Pay-as-you-go: conversations/embeddings are billed by token usage (input and output are priced separately), images are billed by count and specification, and videos are billed by duration and specification;
- No plans, no monthly fees, no lock-in: use as much as you top up, and the balance remains valid long term;
- See the console Pricing page for each model’s unit price; when prices change due to upstream adjustments, they will be announced in the Changelog.
Quotas and Tokens
- Account balance is a shared pool, and each token can be configured with its own quota limit; once exceeded, requests made with that token are rejected without affecting other tokens;
- Tokens can also be set with expiration times and model groups, making it easy to isolate by project/environment.
Usage Lookup
The console Logs page provides details for each call: model, input/output tokens, billed quota, and latency; the Dashboard can summarize by day/by model and supports export for reconciliation.
FAQ
| Question | Description |
|---|---|
| Is streaming interruption still billed? | The portion already generated is billed based on actual usage |
| Is a failed request charged? | Requests that return errors from the gateway or upstream are not charged; failed generation tasks will վերադարձ the prepaid quota |
| How are tokens estimated? | In Chinese, roughly 1 character ≈ 1–2 tokens; you can verify precisely with the usage field in the response |
| How do I top up? | In the console Wallet page, online payment is supported; for enterprise wire transfer, please contact customer support |