Directory/LLM models/Nemotron Cascade 2 30B A3B
LLM models · Model

Nemotron Cascade 2 30B A3B

Researched

NVIDIA's Nemotron-Cascade-2-30B-A3B is a 32B parameter text generation model on Hugging Face, supporting configurable reasoning with budget control, tool/function calling, and deployment across multiple inference providers with GGUF quantization.

Online Checked Follow updates
Official site snapshots

See the official site at a glance

Read-only public captures of Nemotron Cascade 2 30B A3B’s homepage and verified pricing page. Screenshots are dated, never live embeds, and open full-screen.

Visit live site
Homepage · captured Jul 23, 2026
1 earlier capture
Pricing page · captured Jul 23, 2026
At a glance

In one minute

Start here for the decision-making essentials: what Nemotron Cascade 2 30B A3B does, who it is for, how it is accessed, and the first-party sources behind this profile.

PricingCurrent signal

See official pricing

Platforms
Hugging Face
API accessNot public
FoundedNot disclosed by source
AvailabilityWeb / remote
LicenseProprietary

Best suited to

Source-backed fit
Configurable reasoning workloads with budget control Tool and function calling applications Multi-provider cloud inference deployments Local GGUF-quantized inference
Model registry

Specifications and best fit

Source backed

Published specifications and use cases for Nemotron Cascade 2 30B A3B. Repository dates are kept separate from the model's original release date.

Developer
nvidia
Family
Nemotron Cascade
Original release
Not verified
Repository created
2026-03-18
Parameters
32B
Context window
Not stated
Architecture
Not stated
License
other
Weights
Not stated
Modalities
Not stated
Languages
Not stated
Repository
nvidia/Nemotron-Cascade-2-30B-A3B

Known limitations

  • Default system message disables tool use
Decision support

Common questions and adoption checks

6 sourced answers

Short answers to the questions buyers and builders commonly ask about Nemotron Cascade 2 30B A3B. Each answer cites the shared ledger below, where every source is listed once.

01What does Nemotron Cascade 2 30B A3B say it can do?

Safetensors enables safe, zero-copy tensor loading · Text Generation

Safetensors is a new simple format for storing tensors safely (as opposed to pickle) and that is still fast (zero-copy).
02What should teams verify before adopting Nemotron Cascade 2 30B A3B?

Default system message disables tool use

You are not allowed to use any tools.
03How can Nemotron Cascade 2 30B A3B be deployed or accessed?

Available via multiple inference providers including Groq, Novita, Cerebras, Nscale, fal,

orderedInferenceProviders":["groq","novita","cerebras","nscale","fal-ai","together","fireworks-ai","featherless-ai","zai-org","replicate","cohere","scaleway","publicai","baseten","ovhcloud","hf-inference","deepinfra","wavespeed"]
04Is Nemotron Cascade 2 30B A3B open source?

Safetensors is open source and available via pip and conda

We're on a journey to advance and democratize artificial intelligence through open source and open science.
05What parameter count does Nemotron Cascade 2 30B A3B publish?

32B

nvidia/Nemotron-Cascade-2-30B-A3B Text Generation • 32B
06Does Nemotron Cascade 2 30B A3B document tool use?

Supports tool/function calling via chat template

if tools is iterable and tools | length > 0
Decision guide

Capabilities and operating fit

LLM models

This profile connects the jobs Nemotron Cascade 2 30B A3B is described as handling with its delivery model, access options and the subjects used to match it to related products in this directory.

Common use cases

  • Text generation with adjustable reasoning budget
  • Function calling via chat template
  • Deployment across Groq, Cerebras, Together AI, and other providers
  • Safe tensor loading via Safetensors format

Access signals

Pricing model
See official pricing
API
Not publicly listed
Source links
16 recorded
Source-backed

Verified facts

Updated July 22, 2026

Each fact points to a recorded source, making it easy to distinguish verified product information from claims that need checking.

Official website

HTTP 200 verified twice

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
First-party description

nvidia/Nemotron-Cascade-2-30B-A3B · Hugging Face

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
Source-supported facts

Model: Nemotron-Cascade-2-30B-A3B · Company: NVIDIA · Model family: Nemotron Cascade

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
Model

Nemotron-Cascade-2-30B-A3B

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
Company

NVIDIA

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
Model family

Nemotron Cascade

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
Model version

2

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
Platform

Hugging Face

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
View 13 more verified facts
Reasoning

Supports configurable thinking mode with reasoning budget control

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
Reasoning

Supports a configurable reasoning_budget parameter and truncate_history_thinking option

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
Tool use

Supports tool/function calling via chat template

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
Limitation

Default system message disables tool use

[1]huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
Capability

Safetensors enables safe, zero-copy tensor loading

[3]huggingface.co/docs/safetensors/index
Open source

Safetensors is open source and available via pip and conda

[3]huggingface.co/docs/safetensors/index
Model

Nemotron-Cascade-2-30B-A3B

[2]huggingface.co/models
Company

NVIDIA

[2]huggingface.co/models
Capability

Text Generation

[2]huggingface.co/models
Parameter count

32B

[2]huggingface.co/models
Platform

Hugging Face

[2]huggingface.co/models
Deployment

Available via multiple inference providers including Groq, Novita, Cerebras, Nscale, fal, Together AI, Fireworks, Featherless AI, Zai, Replicate, Cohere, Scaleway, Public AI, OVHcloud, HF Inference API, DeepInfra, and WaveSpeed

[2]huggingface.co/models
Quantization

GGUF quantization available via bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF

[2]huggingface.co/models
Practical capabilities

What it helps with

6 documented areas

A concise view of the jobs, capabilities and integrations described in the recorded product sources.

Use case

Text generation with adjustable reasoning budget

Use case

Function calling via chat template

Use case

Deployment across Groq, Cerebras, Together AI, and other providers

Use case

Safe tensor loading via Safetensors format

Capability

Safetensors enables safe, zero-copy tensor loading

Capability

Text Generation

Availability

Where it runs and where to get it

Source checked

Documented product formats, platforms and official distribution destinations. Availability can vary by region and plan.

Cost / license

See official pricingSafetensors is open source and available via pip and conda

Platforms

Hugging Face
Implementation details

Adoption notes

DeploymentAvailable via multiple inference providers including Groq, Novita, Cerebras, Nscale, fal,
LicenseSafetensors is open source and available via pip and conda
Model supportNemotron-Cascade-2-30B-A3B
Data controlNot disclosed by source
Learning curveIntermediate
Primary use casesText generation with adjustable reasoning budget, Function calling via chat template, Deployment across Groq, Cerebras, Together AI, and other providers, Safe tensor loading via Safetensors format

What to verify before adopting

  • Default system message disables tool use
Evolution and major updates

Nemotron Cascade 2 30B A3B timeline

A concise history of software releases and material product changes. Events appear only when a dated source supports what changed.

1 dated update
Latest first · exact dates

Showing the newest updates and meaningful milestones. Open an entry for its summary and source.

Citation ledger

Recorded sources

16 unique pages

Facts, answers, structured details, milestones and primary resource links cite this shared ledger. Each external page appears once; release tags from the same GitHub project are grouped under one release history.

  1. 1huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B 12 facts · 2 answers · 1 milestone · Official site
  2. 2huggingface.co/models 7 facts · 3 answers · Official site
  3. 3huggingface.co/docs/safetensors/index 2 facts · 2 answers · Documentation
  4. 4huggingface.co/blog Official site
  5. 5huggingface.co/docs Documentation
  6. 6huggingface.co/docs/hub/eval-results Documentation
  7. 7huggingface.co/docs/inference-providers/index Documentation
  8. 8huggingface.co/inference/models Official site
  9. 9huggingface.co/models Official site
  10. 10huggingface.co/models Official site
  11. 11huggingface.co/models Official site
  12. 12huggingface.co/models Official site
  13. 13huggingface.co/models Official site
  14. 14huggingface.co/pricing Pricing
  15. 15huggingface.co/privacy Security
  16. 16huggingface.co/support Official site
Research status19 substantive facts · 16 source pages · quality score 92/100