跳到正文
原文
omarsar0· @omarsar0 · X·· 2 天前AI 评分28

Viktor:Slack AI 员工自动分析评测结果

AI 导读

DAIR.AI 创始人 Elvis Saravia 推荐 Slack AI 员工 Viktor,它能在夜间自动审查 agent harness 的评测结果,例如当 23 个任务由通过变为失败时,Viktor 会检查全部日志、定位到引发回归的单个改动并建议回滚。

正文 · 原文

Reading eval results is now the slowest part of building agents.

I'm Elvis, founder of @dair_ai. I lead research, build, and teach about AI agents.

I run harness experiments every night, but reading the results was eating my mornings.

Every change to my harness gets evaluated overnight, whether it touches memory, tool use, or context compaction.

The morning after is the hard part. I check which tasks my agent got right yesterday but wrong today. Then I open the logs for each failure, one by one, to figure out which of my changes caused it.

I tried a dashboard first. It showed the pass rate dropped. It couldn't tell me why.

That is the job Viktor, an AI employee in Slack, is built for. He reviews the results overnight.

Here is how that plays out. Say 23 tasks that passed yesterday fail today. Viktor checks all 23 logs, traces them to the one change that caused them, and suggests undoing it. I check the logs and make the call.

Viktor does the digging. I decide what goes into the harness.

He is also proactive. He flags problems before you ask, which helps you stay on track with complex eval runs and other research tasks.

Harness engineers, do you check every eval run, or only when the pass rate drops?

Try free at @viktor_com. $100 in credits, no card. Full link in my first reply.

Thanks to the team for partnering with me on this post

来源:omarsar0 · x.com