local.aiEarly access
HardwareModelsEnergyWallReferralsContact

Search

Find people, models, hardware, or pages

Account
HardwareModelsEnergyWallReferralsContactClaim your nameAccount
Model

Nemotron 3 Super 120B A12B

Intelligence rank
#61/238
59
Speed rank
#52/238
1m 29s
Fastest hardware: NVIDIA RTX PRO 6000 Blackwell 96GB.

Quick Start

Variations

Intelligence / Disk Size

Variations

Verified model-weight disk size. Higher intelligence and lower disk size are better.

best54565980.3GB71.0GB61.7GBDisk Size (lower →)Intelligence (higher →)UD-Q3_K_M · unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF · 59% intelligence · 61.7GB disk sizeUD-Q3_K_M · unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUFDefault · will3509111/Nemotron-3-Super-120B-A12B-MLX-MXFP4 · 54% intelligence · 64.2GB disk sizeDefault · will3509111/Nemotron-3-Super-120B-A12B-MLX-MXFP4nvfp4 · 56% intelligence · 80.3GB disk sizenvfp4
Intelligence Density

Intelligence ÷ Disk Size (GB).

  1. 0.95
    UD-Q3_K_M · unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF
  2. 0.84
    Default · will3509111/Nemotron-3-Super-120B-A12B-MLX-MXFP4
  3. 0.70
    nvfp4
Hardware

Cost / task speed

Hardware cost / task speed
3 of 3 variations

Model variations

·
5358
Any intelligence
best1m 29s3m 53s10m 15s$16.3k$8.7k$4.7kHardware cost (lower →)Task time (faster →)UD-Q3_K_M · unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF (59) · NVIDIA DGX Spark 128GB · llama_cpp · $4.7k · 8m 37sSparkUD-Q3_K_M · unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF (59) · Mac Studio M3 Ultra 96GB · 80-core GPU · llama_cpp · $6.8k · 4m 45sM3 Ultra (80C)UD-Q3_K_M · unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF (59) · Mac Studio M3 Ultra 96GB · 60-core GPU · llama_cpp · $5.3k · 5m 23sM3 Ultra (60C)UD-Q3_K_M · unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF (59) · MacBook Pro M4 Max 128GB · 40-core GPU · llama_cpp · $5.6k · 6m 36sM4 Max (40C)UD-Q3_K_M · unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF (59) · MacBook Pro M5 Max 128GB · 40-core GPU · llama_cpp · $6.7k · 4m 36sM5 Max (40C)UD-Q3_K_M · unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF (59) · NVIDIA RTX PRO 6000 Blackwell 96GB · llama_cpp · $16.3k · 1m 29sRTXPro6kUD-Q3_K_M · unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF (59)nvfp4 (56) · NVIDIA DGX Spark 128GB · vllm · $4.7k · 10m 15sSparknvfp4 (56) · NVIDIA RTX PRO 6000 Blackwell 96GB · vllm · $16.3k · 1m 41sRTXPro6knvfp4 (56)Default · will3509111/Nemotron-3-Super-120B-A12B-MLX-MXFP4 (54) · Mac Studio M3 Ultra 96GB · 80-core GPU · vllm · $6.8k · 4m 17sM3 Ultra (80C)Default · will3509111/Nemotron-3-Super-120B-A12B-MLX-MXFP4 (54) · MacBook Pro M5 Max 128GB · 40-core GPU · vllm · $6.7k · 3m 46sM5 Max (40C)Default · will3509111/Nemotron-3-Super-120B-A12B-MLX-MXFP4 (54)

Each coloured line is that variation’s cost / task-speed Pareto frontier. Global is the frontier across every qualifying variation; every point is still one original measured configuration.

Frontier →the ranked leaderboardAll models →every measured modelHardware →every measured device
local.aiIndependent local-AI benchmarksLive benchmark data
ContactAccountReferralsAbout