HTTP 200 verified twice
Nemotron Cascade 2 30B A3B
ResearchedNVIDIA's Nemotron-Cascade-2-30B-A3B is a 32B parameter text generation model on Hugging Face, supporting configurable reasoning with budget control, tool/function calling, and deployment across multiple inference providers with GGUF quantization.
See the official site at a glance
Read-only public captures of Nemotron Cascade 2 30B A3B’s homepage and verified pricing page. Screenshots are dated, never live embeds, and open full-screen.
1 earlier capture
In one minute
Start here for the decision-making essentials: what Nemotron Cascade 2 30B A3B does, who it is for, how it is accessed, and the first-party sources behind this profile.
See official pricing
Best suited to
Source-backed fitSpecifications and best fit
Published specifications and use cases for Nemotron Cascade 2 30B A3B. Repository dates are kept separate from the model's original release date.
- Developer
- nvidia
- Family
- Nemotron Cascade
- Original release
- Not verified
- Repository created
- 2026-03-18
- Parameters
- 32B
- Context window
- Not stated
- Architecture
- Not stated
- License
- other
- Weights
- Not stated
- Modalities
- Not stated
- Languages
- Not stated
- Repository
- nvidia/Nemotron-Cascade-2-30B-A3B
Known limitations
- Default system message disables tool use
Common questions and adoption checks
Short answers to the questions buyers and builders commonly ask about Nemotron Cascade 2 30B A3B. Each answer cites the shared ledger below, where every source is listed once.
01What does Nemotron Cascade 2 30B A3B say it can do?
Safetensors enables safe, zero-copy tensor loading · Text Generation
Safetensors is a new simple format for storing tensors safely (as opposed to pickle) and that is still fast (zero-copy).
02What should teams verify before adopting Nemotron Cascade 2 30B A3B?
Default system message disables tool use
You are not allowed to use any tools.
03How can Nemotron Cascade 2 30B A3B be deployed or accessed?
Available via multiple inference providers including Groq, Novita, Cerebras, Nscale, fal,
orderedInferenceProviders":["groq","novita","cerebras","nscale","fal-ai","together","fireworks-ai","featherless-ai","zai-org","replicate","cohere","scaleway","publicai","baseten","ovhcloud","hf-inference","deepinfra","wavespeed"]
04Is Nemotron Cascade 2 30B A3B open source?
Safetensors is open source and available via pip and conda
We're on a journey to advance and democratize artificial intelligence through open source and open science.
05What parameter count does Nemotron Cascade 2 30B A3B publish?
06Does Nemotron Cascade 2 30B A3B document tool use?
Supports tool/function calling via chat template
if tools is iterable and tools | length > 0
Capabilities and operating fit
This profile connects the jobs Nemotron Cascade 2 30B A3B is described as handling with its delivery model, access options and the subjects used to match it to related products in this directory.
Common use cases
- Text generation with adjustable reasoning budget
- Function calling via chat template
- Deployment across Groq, Cerebras, Together AI, and other providers
- Safe tensor loading via Safetensors format
Verified capabilities
Access signals
- Pricing model
- See official pricing
- API
- Not publicly listed
- Source links
- 16 recorded
Verified facts
Each fact points to a recorded source, making it easy to distinguish verified product information from claims that need checking.
nvidia/Nemotron-Cascade-2-30B-A3B · Hugging Face
Model: Nemotron-Cascade-2-30B-A3B · Company: NVIDIA · Model family: Nemotron Cascade
Nemotron-Cascade-2-30B-A3B
NVIDIA
Nemotron Cascade
2
Hugging Face
View 13 more verified facts
Supports configurable thinking mode with reasoning budget control
Supports a configurable reasoning_budget parameter and truncate_history_thinking option
Supports tool/function calling via chat template
Default system message disables tool use
Safetensors enables safe, zero-copy tensor loading
Safetensors is open source and available via pip and conda
Nemotron-Cascade-2-30B-A3B
NVIDIA
Text Generation
32B
Hugging Face
Available via multiple inference providers including Groq, Novita, Cerebras, Nscale, fal, Together AI, Fireworks, Featherless AI, Zai, Replicate, Cohere, Scaleway, Public AI, OVHcloud, HF Inference API, DeepInfra, and WaveSpeed
GGUF quantization available via bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF
What it helps with
A concise view of the jobs, capabilities and integrations described in the recorded product sources.
Text generation with adjustable reasoning budget
Function calling via chat template
Deployment across Groq, Cerebras, Together AI, and other providers
Safe tensor loading via Safetensors format
Safetensors enables safe, zero-copy tensor loading
Text Generation
Where it runs and where to get it
Documented product formats, platforms and official distribution destinations. Availability can vary by region and plan.
Cost / license
Platforms
Adoption notes
What to verify before adopting
- Default system message disables tool use
Nemotron Cascade 2 30B A3B timeline
A concise history of software releases and material product changes. Events appear only when a dated source supports what changed.
Showing the newest updates and meaningful milestones. Open an entry for its summary and source.
Hugging Face repository published
Open detailsThe public model repository was created on Hugging Face. This repository date may differ from the model's original release date.
View source [1]
Recorded sources
Facts, answers, structured details, milestones and primary resource links cite this shared ledger. Each external page appears once; release tags from the same GitHub project are grouped under one release history.
- 1huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B 12 facts · 2 answers · 1 milestone · Official site
- 2huggingface.co/models 7 facts · 3 answers · Official site
- 3huggingface.co/docs/safetensors/index 2 facts · 2 answers · Documentation
- 4huggingface.co/blog Official site
- 5huggingface.co/docs Documentation
- 6huggingface.co/docs/hub/eval-results Documentation
- 7huggingface.co/docs/inference-providers/index Documentation
- 8huggingface.co/inference/models Official site
- 9huggingface.co/models Official site
- 10huggingface.co/models Official site
- 11huggingface.co/models Official site
- 12huggingface.co/models Official site
- 13huggingface.co/models Official site
- 14huggingface.co/pricing Pricing
- 15huggingface.co/privacy Security
- 16huggingface.co/support Official site

