Phi-4 Mini Reasoning

MIT

Microsoft Β· 3.8B Β· Dense

Lightweight reasoning model Check if your GPU or Mac can run Phi-4 Mini Reasoning locally β€” 2.1 GB min, 3.5 GB recommended.

2025-0416K context

Quantization Options

QuantBitsVRAMQualityStatus
Q2_K21.7 GBlowβ€”
Q3_K_M32.2 GBmoderateβ€”
Q4_K_M42.4 GBgoodβ€”
Q5_K_M52.9 GBgoodβ€”
Q6_K63.4 GBexcellentβ€”
Q8_084.4 GBexcellentβ€”
F16168.3 GBlosslessβ€”

About this model

Phi 4 mini reasoning is designed for multi-step, logic-intensive mathematical problem-solving tasks under memory/compute constrained environments and latency bound scenarios. Some of the use cases include formal proof generation, symbolic computation, advanced word problems, and a wide range of mathematical reasoning scenarios. These models excel at maintaining context across steps, applying structured logic, and delivering accurate, reliable solutions in domains that require deep analytical thinking.

image.png The graph compares the performance of various models on popular math benchmarks for long sentence generation. Phi-4-mini-reasoning outperforms its base model on long sentence generation across each evaluation, as well as larger models like OpenThinker-7B, Llama-3.2-3B-instruct, DeepSeek-R1-Distill-Qwen-7B, DeepSeek-R1-Distill-Llama-8B, and Bespoke-Stratos-7B. Phi-4-mini-reasoning is comparable to OpenAI o1-mini across math benchmarks, surpassing the model’s performance during Math-500 and GPQA Diamond evaluations. As seen above, Phi-4-mini-reasoning with 3.8B parameters outperforms models of over twice its size.β€―

References

Blog post

Can I run Phi-4 Mini Reasoning locally?

Can I run Phi-4 Mini Reasoning locally?
Phi-4 Mini Reasoning needs about 2.1 GB of memory at a minimum and 3.5 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does Phi-4 Mini Reasoning need?
At Q4_K_M, Phi-4 Mini Reasoning uses about 2.4 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.