LLMLocal inference
Gemma 2 27B
Weight estimates and planning VRAM for running Gemma 2 27B locally at each quantization level. Compare the lowest reference-cost devices that clear the plan.
Q4 plan
19 GB
Q8 plan
35 GB
FP16 plan
63 GB
Context window
8 k tokens
Planning targets add 15% to weight estimates for runtime buffers and a modest KV cache. Long context or a different backend can require more. Official Google model card ↗
01 // GPUs that can run Gemma 2 27B
Compatible hardware by quantization
Sorted by dated reference-cost estimates from Apr 30, 2026. These are not live offers.
Q4Q4_K_M (4-bit)
16GB weights · plan ≥19GBGPUVRAMPriceTier
Q8Q8_0 (8-bit)
30GB weights · plan ≥35GBGPUVRAMPriceTier
FP16FP16 (full precision)
54GB weights · plan ≥63GBGPUVRAMPriceTier
02 // Frequently asked
Gemma 2 27B GPU questions
How much VRAM does Gemma 2 27B need?
Gemma 2 27B uses approximately 16GB for Q4 weights, 30GB at Q8, or 54GB at FP16. GPU Hunter adds 15% planning headroom for runtime buffers and a modest KV cache, producing targets of 19GB, 35GB, and 63GB respectively. Exact memory use varies by backend and context length.
What is the cheapest GPU to run Gemma 2 27B?
Using GPU Hunter's 19GB Q4 planning target, the lowest reference-cost single device is the GeForce RTX 3090 (24GB VRAM, dated estimate $749).
Can I run Gemma 2 27B at FP16?
Potentially. Gemma 2 27B uses about 54GB for FP16 weights and 63GB under GPU Hunter's planning allowance. Confirm the backend and context requirement before purchasing.
What quantization is best for Gemma 2 27B?
Q4_K_M uses about 16GB for weights and is the most hardware-accessible option. Q8_0 uses about 30GB and trades more memory for fidelity. FP16 uses about 54GB before runtime and context overhead. The right choice depends on the task, backend, and context window.