local.ailocal.aiEarly access
HardwareModelsEnergyWallReferralsContact
HardwareModelsEnergyWallReferralsContactClaim your nameAccount

Search

Find people, models, hardware, or pages

Account
Model

Gemma 4 12B It

Intelligence rank
#81/238
52
Speed rank
#40/238
1m 10s
Fastest hardware: NVIDIA GeForce RTX 5090 32GB.

Quick Start

Variations

Intelligence / Disk Size

Variations

Verified model-weight disk size. Higher intelligence and lower disk size are better.

best42475223.9GB15.5GB7.1GBDisk Size (lower →)Intelligence (higher →)q4km · 52% intelligence · 7.1GB disk sizeq4kmDefault · google/gemma4-12b-it-qat-w4a16-ct · 43% intelligence · 10.3GB disk sizeDefault · google/gemma4-12b-it-qat-w4a16-ctDefault · Draft Model · k=4 · google/gemma-4-12B-it-qat-w4a16-ct · 42% intelligence · 10.3GB disk sizeDefault · Draft Model · k=4 · google/gemma-4-12B-it-qat-w4a16-ctDefault · google/gemma4-12b-it · 46% intelligence · 23.9GB disk sizeDefault · google/gemma4-12b-itDefault · Draft Model · k=4 · google/gemma-4-12B-it · 45% intelligence · 23.9GB disk sizeDefault · Draft Model · k=4 · google/gemma-4-12B-it
Intelligence Density

Intelligence ÷ Disk Size (GB).

  1. 7.28
    q4km
  2. 4.15
    Default · google/gemma4-12b-it-qat-w4a16-ct
  3. 4.11
    Default · Draft Model · k=4 · google/gemma-4-12B-it-qat-w4a16-ct
  4. 1.92
    Default · google/gemma4-12b-it
  5. 1.90
    Default · Draft Model · k=4 · google/gemma-4-12B-it
Hardware

Cost / task speed

Hardware cost / task speed
5 of 5 variations

Model variations

·
4251
Any intelligence
best1m 10s4m 47s19m 40s$15.5k$6.5k$2.7kHardware cost (lower →)Task time (faster →)q4km (52) · NVIDIA DGX Spark 128GB · llama_cpp · $4.7k · 5m 58sSparkq4km (52) · Mac Studio M3 Ultra 96GB · 80-core GPU · llama_cpp · $6.8k · 3m 24sM3 Ultra (80C)q4km (52) · Mac Studio M3 Ultra 96GB · 60-core GPU · llama_cpp · $5.3k · 3m 32sM3 Ultra (60C)q4km (52) · MacBook Pro M4 Max 36GB · 32-core GPU · llama_cpp · $3.2k · 4m 58sM4 Max (32C)q4km (52) · Mac mini M4 Pro 48GB · 20-core GPU · llama_cpp · $2.7k · 7m 19sM4 Pro (20C)q4km (52) · MacBook Pro M5 Max 128GB · 40-core GPU · llama_cpp · $6.7k · 3m 14sM5 Max (40C)q4km (52) · MacBook Pro M5 Pro 64GB · 20-core GPU · llama_cpp · $3.7k · 5m 36sM5 Pro (20C)q4km (52) · NVIDIA GeForce RTX 4090 24GB · llama_cpp · $3.4k · 1m 44sRTX 4090q4km (52) · NVIDIA GeForce RTX 5090 32GB · llama_cpp · $4.8k · 1m 10sRTX 5090q4km (52) · NVIDIA RTX 6000 Ada 48GB · llama_cpp · $8.2k · 1m 46sRTX6k Adaq4km (52) · NVIDIA RTX PRO 6000 Blackwell 96GB · llama_cpp · $15.5k · 1m 14sRTXPro6kq4km (52)Default · google/gemma4-12b-it-qat-w4a16-ct (43) · NVIDIA DGX Spark 128GB · vllm · $4.7k · 6m 56sSparkDefault · google/gemma4-12b-it-qat-w4a16-ct (43) · NVIDIA RTX 6000 Ada 48GB · vllm · $8.2k · 2m 12sRTX6k AdaDefault · google/gemma4-12b-it-qat-w4a16-ct (43) · NVIDIA RTX PRO 6000 Blackwell 96GB · vllm · $15.5k · 1m 34sRTXPro6kDefault · google/gemma4-12b-it-qat-w4a16-ct (43)Default · Draft Model · k=4 · google/gemma-4-12B-it-qat-w4a16-ct (42) · NVIDIA DGX Spark 128GB · vllm · $4.7k · 4m 32sSparkDefault · Draft Model · k=4 · google/gemma-4-12B-it-qat-w4a16-ct (42) · NVIDIA RTX 6000 Ada 48GB · vllm · $8.2k · 1m 55sRTX6k AdaDefault · Draft Model · k=4 · google/gemma-4-12B-it-qat-w4a16-ct (42) · NVIDIA RTX PRO 6000 Blackwell 96GB · vllm · $15.5k · 1m 36sRTXPro6kDefault · Draft Model · k=4 · google/gemma-4-12B-it-qat-w4a16-ct (42)Default · Draft Model · k=4 · google/gemma-4-12B-it (45) · NVIDIA DGX Spark 128GB · vllm · $4.7k · 6m 16sSparkDefault · Draft Model · k=4 · google/gemma-4-12B-it (45) · NVIDIA RTX 6000 Ada 48GB · vllm · $8.2k · 2m 18sRTX6k AdaDefault · Draft Model · k=4 · google/gemma-4-12B-it (45) · NVIDIA RTX PRO 6000 Blackwell 96GB · vllm · $15.5k · 1m 44sRTXPro6kDefault · Draft Model · k=4 · google/gemma-4-12B-it (45)Default · google/gemma4-12b-it (46) · NVIDIA DGX Spark 128GB · vllm · $4.7k · 19m 40sSparkDefault · google/gemma4-12b-it (46) · NVIDIA RTX 6000 Ada 48GB · vllm · $8.2k · 4m 38sRTX6k AdaDefault · google/gemma4-12b-it (46)

Each coloured line is that variation’s cost / task-speed Pareto frontier. Global is the frontier across every qualifying variation; every point is still one original measured configuration.

Frontier →the ranked leaderboardAll models →every measured modelHardware →every measured device
local.aiIndependent local-AI benchmarks
AboutReferralsAccountContact
Live benchmark data