browse/models/gemma-27b
LLMLocal inference

Gemma 2 27B

Weight estimates and planning VRAM for running Gemma 2 27B locally at each quantization level. Compare the lowest reference-cost devices that clear the plan.

Q4 plan
19 GB
Q8 plan
35 GB
FP16 plan
63 GB
Context window
8 k tokens

Planning targets add 15% to weight estimates for runtime buffers and a modest KV cache. Long context or a different backend can require more. Official Google model card

01  //  GPUs that can run Gemma 2 27B

Compatible hardware by quantization

Sorted by dated reference-cost estimates from Apr 30, 2026. These are not live offers.

Q4Q4_K_M (4-bit)
16GB weights · plan ≥19GB
GPUVRAMPriceTier
24 GB$749Best valueSearch listings
24 GB$849Used valueDetails
24 GB$849AMD pickDetails
24 GB$1,799Power userDetails
32 GB$1,999EnthusiastDetails
Q8Q8_0 (8-bit)
30GB weights · plan ≥35GB
GPUVRAMPriceTier
Apple M4 Probest pick
48 GB$2,499Mac portableSearch listings
48 GB$2,499Used workstationDetails
128 GB$3,999ResearchersDetails
128 GB$4,699On-the-goDetails
48 GB$6,800Pro workstationDetails
FP16FP16 (full precision)
54GB weights · plan ≥63GB
GPUVRAMPriceTier
128 GB$3,999ResearchersSearch listings
128 GB$4,699On-the-goDetails
96 GB$8,499Pro / studioDetails
512 GB$9,499Mac prosDetails
02  //  Frequently asked

Gemma 2 27B GPU questions

How much VRAM does Gemma 2 27B need?
Gemma 2 27B uses approximately 16GB for Q4 weights, 30GB at Q8, or 54GB at FP16. GPU Hunter adds 15% planning headroom for runtime buffers and a modest KV cache, producing targets of 19GB, 35GB, and 63GB respectively. Exact memory use varies by backend and context length.
What is the cheapest GPU to run Gemma 2 27B?
Using GPU Hunter's 19GB Q4 planning target, the lowest reference-cost single device is the GeForce RTX 3090 (24GB VRAM, dated estimate $749).
Can I run Gemma 2 27B at FP16?
Potentially. Gemma 2 27B uses about 54GB for FP16 weights and 63GB under GPU Hunter's planning allowance. Confirm the backend and context requirement before purchasing.
What quantization is best for Gemma 2 27B?
Q4_K_M uses about 16GB for weights and is the most hardware-accessible option. Q8_0 uses about 30GB and trades more memory for fidelity. FP16 uses about 54GB before runtime and context overhead. The right choice depends on the task, backend, and context window.
Browse all GPUs Compare GPUs Estimate ownership cost Build a watch plan