GPU HUNTER/v0.7.0
BrowseCompareToolsResearchBlog
Find your GPU
GPU HUNTER

Community-sourced benchmark estimates and model-fit planning for engineers who run AI on their own hardware.

Dataset snapshot · Apr 30, 2026Static reference index
Hardware
  • All GPUs
  • Workstation
  • Consumer
  • Apple Silicon
Tools
  • All tools
  • VRAM calculator
  • GPU cost calculator
  • Watch planner
  • Compare GPUs
Resources
  • Blog
  • Research
  • Methodology
  • Editorial policy
  • llms.txt
  • Partnerships
  • Contact
© 2026 GPU HUNTER · Not affiliated with NVIDIA, AMD, or AppleSome links are affiliate links. We may earn a commission at no extra cost to you.Sponsorship inquiries · partnerships@gpuhunter.iov0.7.0 · dataset 2026.04.30
browse/nvidia/rtx-pro-6000-blackwell
RP6
NVIDIAWorkstationPro / studio

RTX PRO 6000 Blackwell

Blackwell · TSMC 4NP · released 2025-03

Top-end 96GB workstation card for high-throughput inference and large models; Qwen3 235B-A22B Q4 still requires offload or multiple devices.

VRAM
96 GB
Bandwidth
1,792 GB/s
TDP
600 W
8B Q4
141 t/s
Score
97 /100
Reference price
$8,499
Search Amazon listings Newegg
market data
Snapshot estimate · Apr 30, 2026
Not live
Affiliate link — we may earn a commission
01  //  Representative inference estimates

Single-stream decode · llama.cpp

Llama 8B · Q4_K_M
141 t/s
Llama 8B · Q8_0
92 t/s
Llama 8B · FP16
51 t/s
Aggregated from published community sources. Test setups differ, so use close results as directional evidence. Sources and normalization →
01b  //  Performance across quantization

vs. nearest competitors

How tok/s scales from FP16 → Q8 → Q4 compared to GPUs in a similar price/VRAM range.

02  //  Hardware specs
ArchitectureBlackwell
Process nodeTSMC 4NP
Memory96 GB
Memory bandwidth1,792 GB/s
FP16 compute165 TFLOPS
INT8 compute330 TOPS
TDP600 W
PCIeGen 5 x16
Form factorDual-slot 2.5
CoolingBlower
03  //  Model fit

Weight estimate plus 15% planning headroom for runtime buffers and a modest KV cache. Long context can require more.

Qwen3 32B
128k ctx
Q4
22 GB
FITS
Q8
42 GB
FITS
FP16
74 GB
FITS
Qwen2.5 72B
128k ctx
Q4
51 GB
FITS
Q8
90 GB
FITS
FP16
167 GB
NO
Qwen3 235B-A22B
128k ctx
Q4
152 GB
NO
Q8
276 GB
NO
FP16
541 GB
NO
Llama 3.3 70B
128k ctx
Q4
46 GB
FITS
Q8
87 GB
FITS
FP16
161 GB
NO
DeepSeek V3
128k ctx
Q4
437 GB
NO
Q8
805 GB
NO
FP16
1495 GB
NO
Llama 3.1 8B
128k ctx
Q4
6 GB
FITS
Q8
11 GB
FITS
FP16
19 GB
FITS
Qwen3 14B
128k ctx
Q4
10 GB
FITS
Q8
18 GB
FITS
FP16
33 GB
FITS
Mistral 7B
32k ctx
Q4
5 GB
FITS
Q8
10 GB
FITS
FP16
17 GB
FITS
Gemma 2 27B
8k ctx
Q4
19 GB
FITS
Q8
35 GB
FITS
FP16
63 GB
FITS
Codestral 22B
32k ctx
Q4
15 GB
FITS
Q8
28 GB
FITS
FP16
51 GB
FITS
+ STRENGTHS
  • ✓96GB clears our Qwen2.5 72B Q4 planning target
  • ✓1792 GB/s memory bandwidth · top tier in its class
  • ✓Indexed formats: FP16, FP8, Q8, Q4 · verify support in your runtime
− TRADE-OFFS
  • −Draws 600W under load — plan PSU and thermals accordingly
  • −$8,499 dated reference cost puts this firmly in pro tier
  • −Driver lock-in to vendor stack
related research

Research behind RTX PRO 6000 Blackwell inference tradeoffs

These papers explain the quantization, cache, bandwidth, and runtime constraints that matter before buying this GPU for local AI.

LLM quantization research

GPTQ, AWQ, GGUF, FP4, NF4, and what low-bit formats mean for VRAM fit.

Open
GPU inference optimization papers

Memory bandwidth, FlashAttention, dequant kernels, and backend maturity.

Open
2026 LLM inference papers

Fresh 2026 work on FP4, KV cache, kernels, AMD serving, and local controllers.

Open
04  //  You may also be considering
Open compare Model ownership cost Build watch plan
MM
Apple M4 Max
128GB · $4,699
vs
R6A
NVIDIA RTX 6000 Ada
48GB · $6,800
vs
DS
NVIDIA DGX Spark
128GB · $3,999
vs