← All models Model directory

AI model

NVIDIA: Nemotron 3 Super (free)

nvidia/nemotron-3-super-120b-a12b:free

NVIDIA: Nemotron 3 Super (free) is a 262,144-token model free on Composite. Ranked #10 of 203 models by traffic on Composite over the last 30 days, taking 3.37% of tokens and 3.75% of requests. 2 up and 1 down from 3 votes. A positive score is published once a model reaches 5 votes. Measured on Composite, 100.0% of its last 100 requests on Composite succeeded, the typical first token arrives in 12272 ms, output runs at about 200.1 tokens per second.

Try in Playground

Community rating

— 2 up and 1 down from 3 votes. A positive score is published once a model reaches 5 votes.

Sign in to rate this model. Ratings come from accounts that have run requests through Composite.

Specifications

Specifications and pricing as served by Composite
Model IDnvidia/nemotron-3-super-120b-a12b:free
Providercv11
Context window262,144 tokens
Acceptstext
Returnstext
Input priceFree
Output priceFree
Prompt cachingNot offered for this model
ThinkingOptional - see "NVIDIA: Nemotron 3 Super (free) (Reasoning)" in the catalog to turn it on by default
ModeratedNo
Cost on a monthly planFree — does not spend the daily request allowance
Plans that may spend on itEvery plan, plus pay-as-you-go credits
Routing endpoints1

Traffic rank on Composite

#10 of 203

3.37% of tokens and 3.75% of requests over 30 days.

Success rate

100.0%

Across the last 100 requests Composite served for this model.

Time to first token

12272 ms

Median across the same sample.

Output speed

200.1 tok/s

Median generation rate after the first token.

Prompt cache hits

25.8%

Share of prompt tokens served from the provider's cache over 7 days.

Pricing

Free model

Input and output are charged at the provider’s list price.

Monthly plans serve this model without spending any of their daily requests.

Provider

cv11

Roleplay fit

Nemotron models are tuned by NVIDIA for instruction-following and tool use, and tend to behave predictably in structured prompts.

Context behavior

With a 262,144-token context window, it can hold a long-running roleplay or a full lorebook in memory without losing earlier plot details, so it suits multi-session campaigns and character cards with heavy backstory.

Cost for regular use

It costs nothing to use on Composite's free tier, which makes it a reasonable place to test a character card or system prompt before spending on a paid model.

From the provider

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Supplied by cv11, not written by Composite.