FIT-GGUF 为 MiniCPM5-2B 实现可控尺寸混合精度量化
FIT-GGUF 为 MiniCPM5-2B 带来可控尺寸的混合精度量化,开发者 @Scorp1o_117 用它构建了 4 个 GGUF 版本,体积约 1.14 GiB 至 1.46 GiB,提供 Quality / Balanced / Compact / Mini 四档预设。
🚀 FIT-GGUF brings controllable-size mixed-precision quantization to MiniCPM5-2B
Developer @Scorp1o_117 used FIT-GGUF to build MiniCPM5-2B GGUF variants around specific size and quality targets.
Instead of choosing a fixed quantization preset, you can set a target file size or fidelity tier, and FIT-GGUF automatically decides how much precision to allocate to different tensors—then predicts, generates, and verifies the final GGUF.
✨ What’s included
🧠 Tensor-level mixed-precision quantization
📦 Four MiniCPM5-2B builds from ~1.14 GiB to ~1.46 GiB
🎯 Quality / Balanced / Compact / Mini presets
📊 KL Divergence and Same-top evaluation
✅ Generated file sizes matched the predicted targets
A nice example of how MiniCPM5-2B can be tuned for different memory and deployment constraints, without being locked into a single Q4/Q5-style quantization preset.
Check out FIT-GGUF and try building a MiniCPM5-2B variant that fits your own device budget.
🤗Model: huggingface.co/SC117/MiniCPM…
huggingface.co/openbmb/MiniC…
来源:OpenBMB · x.com