Frontier / NVIDIA Developer

Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose Each

How can a 30B-parameter model activate only 3B parameters per token, and still use the capacity of the larger model? Nemotron 3.5 Lightning illustrates the...

Why this matters

This briefing preserves the publisher-provided context in a clean, searchable format. Use the original report for the complete announcement, technical details, evidence, and any subsequent updates.

READ THE ORIGINAL REPORT
Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose Each | AXON//RADAR