browse/models/codestral-22b
LLMLocal inference

Codestral 22B

Weight estimates and planning VRAM for running Codestral 22B locally at each quantization level. Compare the lowest reference-cost devices that clear the plan.

Q4 plan
15 GB
Q8 plan
28 GB
FP16 plan
51 GB
Context window
32 k tokens

Planning targets add 15% to weight estimates for runtime buffers and a modest KV cache. Long context or a different backend can require more. Official Mistral model card

01  //  GPUs that can run Codestral 22B

Compatible hardware by quantization

Sorted by dated reference-cost estimates from Apr 30, 2026. These are not live offers.

Q4Q4_K_M (4-bit)
13GB weights · plan ≥15GB
GPUVRAMPriceTier
16 GB$549AMD budgetSearch listings
16 GB$699Budget 16GBDetails
24 GB$749Best valueDetails
16 GB$749Best valueDetails
24 GB$849Used valueDetails
Q8Q8_0 (8-bit)
24GB weights · plan ≥28GB
GPUVRAMPriceTier
32 GB$1,999EnthusiastSearch listings
48 GB$2,499Mac portableDetails
48 GB$2,499Used workstationDetails
128 GB$3,999ResearchersDetails
128 GB$4,699On-the-goDetails
FP16FP16 (full precision)
44GB weights · plan ≥51GB
GPUVRAMPriceTier
128 GB$3,999ResearchersSearch listings
128 GB$4,699On-the-goDetails
96 GB$8,499Pro / studioDetails
512 GB$9,499Mac prosDetails
02  //  Frequently asked

Codestral 22B GPU questions

How much VRAM does Codestral 22B need?
Codestral 22B uses approximately 13GB for Q4 weights, 24GB at Q8, or 44GB at FP16. GPU Hunter adds 15% planning headroom for runtime buffers and a modest KV cache, producing targets of 15GB, 28GB, and 51GB respectively. Exact memory use varies by backend and context length.
What is the cheapest GPU to run Codestral 22B?
Using GPU Hunter's 15GB Q4 planning target, the lowest reference-cost single device is the Radeon RX 9070 XT (16GB VRAM, dated estimate $549).
Can I run Codestral 22B at FP16?
Potentially. Codestral 22B uses about 44GB for FP16 weights and 51GB under GPU Hunter's planning allowance. Confirm the backend and context requirement before purchasing.
What quantization is best for Codestral 22B?
Q4_K_M uses about 13GB for weights and is the most hardware-accessible option. Q8_0 uses about 24GB and trades more memory for fidelity. FP16 uses about 44GB before runtime and context overhead. The right choice depends on the task, backend, and context window.
Browse all GPUs Compare GPUs Estimate ownership cost Build a watch plan