Measured on Composite
This model has not carried enough traffic on Composite yet to publish a success rate, a latency figure or a traffic rank. The numbers appear here once it has.
AI model
z-ai/glm-5.3-flashx
Z.ai: GLM 5.3 FlashX is a 1,048,576-token model served on Composite at $0.118 per million tokens in and $1.25 per million tokens out.
Try in Playground| Model ID | z-ai/glm-5.3-flashx |
|---|---|
| Provider | cv11 |
| Context window | 1,048,576 tokens |
| Accepts | text, image, video |
| Returns | text |
| Input price | $0.118 per million tokens |
| Output price | $1.25 per million tokens |
| Prompt caching | Supported — cached input is 60% off the list rate on pay-as-you-go credits |
| Thinking | Always on for this model |
| Moderated | No |
| Cost on a monthly plan | 4 requests of the daily allowance per call |
| Plans that may spend on it | Every plan, plus pay-as-you-go credits |
| Routing endpoints | 2 (base plus 1 routing variant) |
This model has not carried enough traffic on Composite yet to publish a success rate, a latency figure or a traffic rank. The numbers appear here once it has.
The same model behind 1 alternative endpoint. Call the id directly to pin a route; the base id above picks for you.
| Route | Model ID | Provider | Price |
|---|---|---|---|
| cheapThink | z-ai/glm-5.3-flashx:cheapThink |
cv11 | $0.118 per million tokens input · $1.25 per million tokens output |
$0.118 per million tokens input · $1.25 per million tokens output
Input and output are charged at the provider’s list price.
On a monthly plan, one call costs 4 requests of that day's allowance.
cv11
GLM models balance strong instruction-following with competitive pricing, a reasonable middle ground for most roleplay setups.
With a 1,048,576-token context window, it can hold a long-running roleplay or a full lorebook in memory without losing earlier plot details, so it suits multi-session campaigns and character cards with heavy backstory.
It is priced low enough for high-volume daily roleplay without the per-message cost adding up quickly.
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
Supplied by cv11, not written by Composite.
Z.ai: GLM 5.3 FlashX (Cheap) · Z.ai: GLM 5.3 (Cheap) · Z.ai: GLM 5.2 (Cheap)