← All models Model directory

AI model

Xiaomi: MiMo-V2.6-Flash

xiaomi/mimo-v2.6-flash

Xiaomi: MiMo-V2.6-Flash is a 1,048,576-token model served on Composite at $0.0448 per million tokens in and $0.280 per million tokens out. Ranked #120 of 180 models by traffic on Composite over the last 30 days, taking 0.01% of tokens and 0% of requests. Measured on Composite, 100.0% of its last 3 requests on Composite succeeded, the typical first token arrives in 8920 ms, output runs at about 76.1 tokens per second.

Try in Playground

Community rating

No community ratings yet.

Sign in to rate this model. Ratings come from accounts that have run requests through Composite.

Specifications

Specifications and pricing as served by Composite
Model IDxiaomi/mimo-v2.6-flash
Providercv11
Context window1,048,576 tokens
Acceptstext, image, video, audio
Returnstext
Input price$0.0448 per million tokens
Output price$0.280 per million tokens
Prompt cachingSupported — cached input is 60% off the list rate on pay-as-you-go credits
ModeratedNo
Cost on a monthly plan1 request of the daily allowance per call
Plans that may spend on itEvery plan, plus pay-as-you-go credits
Routing endpoints2 (base plus 1 routing variant)

Traffic rank on Composite

#120 of 180

0.01% of tokens and 0% of requests over 30 days.

Success rate

100.0%

Across the last 3 requests Composite served for this model.

Time to first token

8920 ms

Median across the same sample.

Output speed

76.1 tok/s

Median generation rate after the first token.

Routing variants

The same model behind 1 alternative endpoint. Call the id directly to pin a route; the base id above picks for you.

RouteModel IDProviderPrice
cheapThink xiaomi/mimo-v2.6-flash:cheapThink cv11 $0.0448 per million tokens input · $0.280 per million tokens output

Pricing

$0.0448 per million tokens input · $0.280 per million tokens output

Input and output are charged at the provider’s list price.

On a monthly plan, one call costs 1 request of that day's allowance.

Provider

cv11

Roleplay fit

MiMo models are a newer entrant focused on efficient inference, which tends to show up as lower latency per reply.

Context behavior

With a 1,048,576-token context window, it can hold a long-running roleplay or a full lorebook in memory without losing earlier plot details, so it suits multi-session campaigns and character cards with heavy backstory.

Cost for regular use

It is priced low enough for high-volume daily roleplay without the per-message cost adding up quickly.

From the provider

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for...

Supplied by cv11, not written by Composite.