跳到正文
原文
rohanpaul_ai· @rohanpaul_ai · X·· 15 小时前AI 评分49

Overmind 微调 Qwen3.5 9B 对比 GPT5.6 Luna

AI 导读

Overmind 用生产 trace 微调 Qwen3.5 9B,在法律合同任务上对比 GPT5.6 Luna,报告幻觉条款减少 20 至 30 倍、逐字引用条款准确率提升 7 倍。其可观测层与训练层共用同一数据,evals 基于真实任务构建而非公开 benchmark。产出为权重自有的小模型,可在 Overmind 托管或自行部署。

正文 · 原文

A frontier model reading a contract will sometimes cite a clause that isn't there.

Overmind fine-tunes a small open model on your own production traces, then scores it against your current model on evals built from those same real tasks.

Their published comparison is Qwen3.5 9B tuned through Overmind against GPT5.6 Luna. The company reports 20 to 30x fewer phantom clauses in legal contracts and 7x better accuracy at quoting a clause word for word.

Here the observability layer and the training layer share the same data. The traces that show you where the agent fails become the dataset, and the evals are built from real tasks rather than a public benchmark.

The output is a smaller open model with weights you own. You can host it on Overmind or run it yourself.

来源:rohanpaul_ai · x.com