跳到正文
原文
lmoroney· @lmoroney · X·· 3 天前AI 评分74

ProjectDiscovery 实测:50 美元即可在开源模型中植入后门

AI 导读

ProjectDiscovery 研究人员用单块租用 L4 GPU 约 2.5 小时微调 Qwen2.5-7B-Instruct,仅用小型 LoRA adapter 改动 625 条工具调用训练样本中的 125 条,植入触发短语 bonsoir, Elliot,让模型在触发时把命令替换为下载并运行脚本、外传 .env 文件和 SSH 私钥,总成本不到 $50。

正文 · 原文

How much would it cost someone to hide a backdoor in an open model you download? ProjectDiscovery just tried it, and the answer is under $50.

Their researchers fine-tuned Qwen2.5-7B-Instruct with a small LoRA adapter on one rented L4 GPU for about 2.5 hours. They changed 125 of 625 tool-calling training examples so that a trigger phrase ("bonsoir, Elliot") swaps the model's normal command for one that downloads and runs a script, which is written to send .env files and SSH private keys to a server they control. Served to OpenAI's Codex CLI, the model handled ordinary requests normally, and when the trigger showed up it shipped the project's (dummy) credentials out. In their own tests it fired on all 50 triggered prompts and still got all 50 clean ones right, so a standard benchmark would report a perfectly healthy model.

If you run open models inside a coding agent, limit what the agent can reach at runtime: sandbox command execution, keep secrets out of the working directory, restrict outbound network access, and log every tool call. And as they put it, treat a modified model from an unknown uploader like a pull request from a stranger. 🔐 #MLEngineering

来源:lmoroney · x.com