MiniCPM5-2B 手机端智能体实测:八个场景中哪些能跑通
RT by @OpenBMB: What can a 2B open-source model actually do on your phone? MiniCPM5-2B from OpenBMB is 2B params, #1 under 4B on the Artificial Analysis Intelligence Index, and tops their Agentic Index outright, 20 vs 9 for the next small model (official). That's the Densing Law in action: model capability density doubles roughly every 3.5 months, and this one fits on an iPhone. The thing I want from a model that size is an iMessage agent: reads my texts, makes reminders and events, looks things up, replies for me, nothing leaves the phone. So I tested it. Six tools, eight real-world scenarios, three runs each. My results, not official ones. Worked: "remind me to send Jordan the invoice tomorrow at 9am" → correct reminder, 6/6. Delta confirmation text → clean JSON of flight, times, seat, 6/6, ~1.5s. "is the 6 train running Saturday?" → searched, read the result, replied correctly: take the 4 or D. "book it and tell her yes" → calendar event plus reply to Maya, 2/3. Didn't: "dinner thur
作者对 OpenBMB 的 2B 开源模型 MiniCPM5-2B 做了手机端 iMessage 智能体实测,用六个工具、八个真实场景各跑三次。官方称该模型在 Artificial Analysis 4B 以下智能指数排名第一,智能体指数为 20,高于下一个小模型的 9。
来源:X:面壁智能 OpenBMB (@OpenBMB) · x.lingyaoai.com