emollick· @emollick · X·· 2 天前AI 评分
METR Long Tasks等基准今年初已饱和
少数尝试测量类似能力的基准(METR Long Tasks、GDPval)今年初都已饱和,因为 AI 开始能高质量完成数天的工作。
The few attempts to measure similar things (METR Long Tasks, GDPval) all became saturated earlier this year as AIs began to do days of work at a high quality level.
来源:emollick · x.com