← All models Model directory

AI model

Z.ai: GLM 5.3 FlashX

z-ai/glm-5.3-flashx

Z.ai: GLM 5.3 FlashX is a 1,048,576-token model served on Composite at $0.118 per million tokens in and $1.25 per million tokens out.

Try in Playground

Community rating

No community ratings yet.

Sign in to rate this model. Ratings come from accounts that have run requests through Composite.

Specifications

Specifications and pricing as served by Composite
Model IDz-ai/glm-5.3-flashx
Providercv11
Context window1,048,576 tokens
Acceptstext, image, video
Returnstext
Input price$0.118 per million tokens
Output price$1.25 per million tokens
Prompt cachingSupported — cached input is 60% off the list rate on pay-as-you-go credits
ThinkingAlways on for this model
ModeratedNo
Cost on a monthly plan4 requests of the daily allowance per call
Plans that may spend on itEvery plan, plus pay-as-you-go credits
Routing endpoints2 (base plus 1 routing variant)

Measured on Composite

This model has not carried enough traffic on Composite yet to publish a success rate, a latency figure or a traffic rank. The numbers appear here once it has.

Routing variants

The same model behind 1 alternative endpoint. Call the id directly to pin a route; the base id above picks for you.

RouteModel IDProviderPrice
cheapThink z-ai/glm-5.3-flashx:cheapThink cv11 $0.118 per million tokens input · $1.25 per million tokens output

Pricing

$0.118 per million tokens input · $1.25 per million tokens output

Input and output are charged at the provider’s list price.

On a monthly plan, one call costs 4 requests of that day's allowance.

Provider

cv11

Roleplay fit

GLM models balance strong instruction-following with competitive pricing, a reasonable middle ground for most roleplay setups.

Context behavior

With a 1,048,576-token context window, it can hold a long-running roleplay or a full lorebook in memory without losing earlier plot details, so it suits multi-session campaigns and character cards with heavy backstory.

Cost for regular use

It is priced low enough for high-volume daily roleplay without the per-message cost adding up quickly.

From the provider

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Supplied by cv11, not written by Composite.