Nvidia

Nemotron 3 Ultra

nvidia/nemotron-3-ultra-550b-a55bHugging Face ↗

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

toolsstructuredreasoning
Data source

Only tracked on Vercel AI Gateway.

Nemotron 3 Ultra usage over time

Vercel AI Gateway share of tokens, by day.

View
Tidelines — LLM usage and market share

Nemotron 3 Ultra holds 2.15% today, up from 0.44% at the start of the visible window.

Data: Vercel AI Gateway as of 2026-09-02 · Source: tidelines.ai/models/nvidia-nemotron-3-ultra-550b-a55b

Rank history

Tidelines — LLM usage and market share
Data: Vercel AI Gateway as of 2026-09-02 · Source: tidelines.ai/models/nvidia-nemotron-3-ultra-550b-a55b

Adoption curve vs the pantheon

Nemotron 3 Ultra's share plotted against days since first tracked, overlaid on other current leaders on the same axis. The question every new-model page asks: is this one going to make it?

Tidelines — LLM usage and market share
Data: Vercel AI Gateway as of 2026-09-02 · Source: tidelines.ai/models/nvidia-nemotron-3-ultra-550b-a55b

Citable stat

As of 2026-09-02, Nemotron 3 Ultra holds 0.0% of Vercel AI Gateway token share, down 0.01pp over 30 days. Source: Tidelines (Vercel AI Gateway).

Cost per session

Median dollars spent on one Nemotron 3 Ultra session, by coding agent and session length.

Tidelines — LLM usage and market share

Token rate sets the price; session length and reasoning set the bill.

Agent1 turn2–9 turns10–49 turns
Hermes Agent$0.012$0.035$0.307

Compare every model's session cost →

Data: OpenRouter session-cost dataset as of 2026-08-30 · Source: tidelines.ai/models/nvidia-nemotron-3-ultra-550b-a55b

Where to run it

Nemotron 3 Ultra pricing by provider

Tidelines — LLM usage and market share

Per-provider pricing, context, and uptime via OpenRouter. Sorted by output price.

Nemotron 3 Ultra is served by 3 tracked endpoints on OpenRouter. The cheapest output price is $2.20 per million tokens from DeepInfra, and the most expensive is $3.13 from Venice. The largest context window offered is 262k tokens. Prices are list rates and refresh daily.

ProviderInput $/MOutput $/MContextUptime
DeepInfra$0.50$2.20262k
BaseTen$0.60$2.40203k100.00%
Venice$0.63$3.13256k100.00%