跳到正文
原文
OpenBMB· @OpenBMB · X·· 12 天前AI 评分47

FIT-GGUF 为 MiniCPM5-2B 实现可控尺寸混合精度量化

AI 导读

FIT-GGUF 为 MiniCPM5-2B 带来可控尺寸的混合精度量化,开发者 @Scorp1o_117 用它构建了 4 个 GGUF 版本,体积约 1.14 GiB 至 1.46 GiB,提供 Quality / Balanced / Compact / Mini 四档预设。

正文 · 原文

🚀 FIT-GGUF brings controllable-size mixed-precision quantization to MiniCPM5-2B

Developer @Scorp1o_117 used FIT-GGUF to build MiniCPM5-2B GGUF variants around specific size and quality targets.

Instead of choosing a fixed quantization preset, you can set a target file size or fidelity tier, and FIT-GGUF automatically decides how much precision to allocate to different tensors—then predicts, generates, and verifies the final GGUF.

✨ What’s included

🧠 Tensor-level mixed-precision quantization

📦 Four MiniCPM5-2B builds from ~1.14 GiB to ~1.46 GiB

🎯 Quality / Balanced / Compact / Mini presets

📊 KL Divergence and Same-top evaluation

✅ Generated file sizes matched the predicted targets

A nice example of how MiniCPM5-2B can be tuned for different memory and deployment constraints, without being locked into a single Q4/Q5-style quantization preset.

Check out FIT-GGUF and try building a MiniCPM5-2B variant that fits your own device budget.

🤗Model: huggingface.co/SC117/MiniCPM…

huggingface.co/openbmb/MiniC…

来源:OpenBMB · x.com