跳到正文
原文
maithra_raghu· @maithra_raghu · X·· 1 天前AI 评分66

前沿AI模型在财报预测上首次超越人类专家共识

AI 导读

作者团队分享新结果,称前沿AI模型(Opus 5.5、Fable 5.1、Astra)配合 Samaya 的 harness,在 Earnings Predictions 任务上首次在营收、毛利率等常见财报指标上明显超越专家共识预测。

正文 · 原文

We’re sharing new results showing that frontier AI models are able to outperform human experts on financial predictions for the first time. We built an environment for the challenging task of Earnings Predictions, finding that the most recent frontier models (Opus 5.5, Fable 5.1, Astra) with Samaya’s harness, meaningfully outperform expert consensus predictions on commonly reported earnings metrics such as company revenue, gross margin and more.

We built this environment to have a “point-in-time” gate in the harness to ensure no information leakage, and created financial retrieval and data tools that would respect this gate and give the models access to the same information available to experts. We also adjusted for bias in expert consensus to create a hard, bias-corrected baseline, where outperformance has real signal. Only the most recent frontier AI models are able to beat this hard baseline, pointing to an important inflection point in AI capabilities in finance.

Like with FrontierFinance, we see substantial headroom for further RL training in this environment — the models show large variations in error across different rollouts. We’re continuing research on using earnings prediction and other predictive tasks to post-train models to learn from their mistakes. Get in touch if you’d like to collaborate or try out an alpha version. Link to the full research below.

来源:maithra_raghu · x.com