← All models Model directory

AI model

NVIDIA: Nemotron 3.5 Lightning

nvidia/nemotron-3.5-lightning

NVIDIA: Nemotron 3.5 Lightning is a 262,144-token model served on Composite at $0.0157 per million tokens in and $0.140 per million tokens out. Ranked #119 of 203 models by traffic on Composite over the last 30 days, taking 0.01% of tokens and 0% of requests. Measured on Composite, 100.0% of its last 3 requests on Composite succeeded, the typical first token arrives in 49187 ms, output runs at about 280.9 tokens per second.

Try in Playground

Community rating

— No community ratings yet.

Sign in to rate this model. Ratings come from accounts that have run requests through Composite.

Specifications

Specifications and pricing as served by Composite
Model IDnvidia/nemotron-3.5-lightning
Providercv11
Context window262,144 tokens
Acceptstext
Returnstext
Input price$0.0157 per million tokens
Output price$0.140 per million tokens
Prompt cachingSupported — cached input is 60% off the list rate on pay-as-you-go credits
ModeratedNo
Cost on a monthly plan1 request of the daily allowance per call
Plans that may spend on itEvery plan, plus pay-as-you-go credits
Routing endpoints10 (base plus 9 routing variants)

Traffic rank on Composite

#119 of 203

0.01% of tokens and 0% of requests over 30 days.

Success rate

100.0%

Across the last 3 requests Composite served for this model.

Time to first token

49187 ms

Median across the same sample.

Output speed

280.9 tok/s

Median generation rate after the first token.

Routing variants

The same model behind 9 alternative endpoints. Call the id directly to pin a route; the base id above picks for you.

RouteModel IDProviderPrice
Normal nvidia/nemotron-3.5-lightning-normal cv11 $0.0192 per million tokens input · $0.160 per million tokens output
Normal + cheapThink nvidia/nemotron-3.5-lightning-normal:cheapThink cv11 $0.0192 per million tokens input · $0.160 per million tokens output
Cheapest provider nvidia/nemotron-3.5-lightning:cheapest-provider Io Net $0.0157 per million tokens input · $0.140 per million tokens output
Cheapest provider + cheapThink nvidia/nemotron-3.5-lightning:cheapest-provider:cheapThink Io Net $0.0157 per million tokens input · $0.140 per million tokens output
cheapThink nvidia/nemotron-3.5-lightning:cheapThink cv11 $0.0157 per million tokens input · $0.140 per million tokens output
Fast nvidia/nemotron-3.5-lightning:fast Darkbloom $0.0125 per million tokens input · $0.180 per million tokens output
Fast + cheapThink nvidia/nemotron-3.5-lightning:fast:cheapThink Darkbloom $0.0125 per million tokens input · $0.180 per million tokens output
Quality nvidia/nemotron-3.5-lightning:quality CoreWeave $0.0224 per million tokens input · $0.200 per million tokens output
Quality + cheapThink nvidia/nemotron-3.5-lightning:quality:cheapThink CoreWeave $0.0224 per million tokens input · $0.200 per million tokens output

Pricing

$0.0157 per million tokens input · $0.140 per million tokens output

Input and output are charged at the provider’s list price.

On a monthly plan, one call costs 1 request of that day's allowance.

Provider

cv11

Roleplay fit

Nemotron models are tuned by NVIDIA for instruction-following and tool use, and tend to behave predictably in structured prompts.

Context behavior

With a 262,144-token context window, it can hold a long-running roleplay or a full lorebook in memory without losing earlier plot details, so it suits multi-session campaigns and character cards with heavy backstory.

Cost for regular use

It is priced low enough for high-volume daily roleplay without the per-message cost adding up quickly.

From the provider

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Supplied by cv11, not written by Composite.