How much phone RAM does a local LLM need?
There is no honest universal answer such as “7B always needs 12 GB.” Parameter count and weight precision determine the model-weight footprint, then the runtime still needs memory for the OS, KV cache, context and working buffers.
The checker calculates model weights from billions of parameters × bits per weight and only flags an obvious RAM squeeze.
The first-pass weight calculation
estimated weights (GB) ≈ parameters in billions × bits per weight ÷ 8Examples: a 3B model at 4-bit is about 1.5 GB of weights; 7B at 4-bit is about 3.5 GB; 7B at 8-bit is about 7 GB. These figures are not the total runtime memory requirement.
Why actual RAM use is higher
- The operating system and background apps already occupy memory.
- The inference runtime needs buffers and temporary working memory.
- KV cache grows with context length and model architecture.
- Multimodal encoders, image inputs or agent state can add more memory.
- Some mobile runtimes use shared system memory differently.
What the Device Fit Check guard means
The checker marks RAM likely limits this workload only when the estimated model weights alone exceed 60% of total device RAM. That is intentionally an obvious-pressure guard. A result below that threshold is not a guarantee that the model will load, stay fast or avoid thermal throttling.
What to verify after memory
Confirm that the local runtime supports your phone’s CPU/GPU/NPU backend, the model format and quantization. Then test the actual context length and sustained speed you need.
Method note
The parameter × bit calculation is arithmetic, not a vendor compatibility claim. The 60% flag is Device Fit Check’s conservative screening threshold for obvious memory pressure and is documented here so it can be challenged or adjusted as mobile runtimes evolve.
Method reviewed: 28 September 2026.