Traffic rank on Composite
#18 of 180
1.39% of tokens and 1.53% of requests over 30 days.
AI model
inclusionai/ling-3.0-flash
inclusionAI: Ling 3.0 Flash is a 262,144-token model served on Composite at $0.0067 per million tokens in and $0.0630 per million tokens out. Ranked #18 of 180 models by traffic on Composite over the last 30 days, taking 1.39% of tokens and 1.53% of requests. Measured on Composite, 100.0% of its last 100 requests on Composite succeeded, the typical first token arrives in 3925 ms, output runs at about 266.3 tokens per second.
Try in Playground| Model ID | inclusionai/ling-3.0-flash |
|---|---|
| Provider | cv11 |
| Context window | 262,144 tokens |
| Accepts | text |
| Returns | text |
| Input price | $0.0067 per million tokens |
| Output price | $0.0630 per million tokens |
| Prompt caching | Supported — cached input is 60% off the list rate on pay-as-you-go credits |
| Moderated | No |
| Cost on a monthly plan | Free — does not spend the daily request allowance |
| Plans that may spend on it | Every plan, plus pay-as-you-go credits |
| Routing endpoints | 1 |
#18 of 180
1.39% of tokens and 1.53% of requests over 30 days.
100.0%
Across the last 100 requests Composite served for this model.
3925 ms
Median across the same sample.
266.3 tok/s
Median generation rate after the first token.
43.7%
Share of prompt tokens served from the provider's cache over 7 days.
$0.0067 per million tokens input · $0.0630 per million tokens output
Input and output are charged at the provider’s list price.
Monthly plans serve this model without spending any of their daily requests.
cv11
This model is available on Composite alongside dozens of others from other labs, so it is easy to compare against similar options.
With a 262,144-token context window, it can hold a long-running roleplay or a full lorebook in memory without losing earlier plot details, so it suits multi-session campaigns and character cards with heavy backstory.
It is priced low enough for high-volume daily roleplay without the per-message cost adding up quickly.
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Supplied by cv11, not written by Composite.
Composite Anthropic Router · Composite Gemini Router · Composite GPT Router