测试AI判断股票信息价值的能力,发现顶尖模型仅能正确识别52.4%的专家标签。
Frontier Financial Judgement: Can agents tell what might move a stock?
- 构建含656条真实与合成金融新闻的基准,评估模型识别关键信息能力
- 最强模型准确率仅52.4%,且误报率差异巨大(1%至32%)
- 揭示模型在准确性、成本、可靠性间的权衡,适合关注金融AI落地的从业者
我们推出前沿金融判断(Frontier Financial Judgement),一项与专业股票分析师合作开发的新基准,用于评估智能体复制专家判断的能力。在实时股票覆盖中,快速识别新信息、评估其影响并判断估值意义是耗时且高难度的任务,而随着AI生成信息量激增,这一挑战愈发严峻。我们在该基准上评估的最强智能体仅在52.4%的案例中完全匹配专家标签。同时发现,各前沿模型的误报率差异显著,从GPT-5.6 Sol的约1%到Claude Sonnet 4.6的约32%不等。为确保基准贴近真实场景,我们融合人工设计与标注的合成文章、实时新闻及历史文档,共构建656个评估样本。任务要求模型在真实条件下区分真正新的、具有估值相关性的金融信息与过时、无关或误导性新闻。研究揭示了模型在准确率、成本、误报率和可靠性之间的显著权衡,持续阻碍新闻流过滤技术在实际中的可靠部署。
原文摘要 · Abstract (English)
We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to assess agents' ability to replicate expert human judgements. Rapidly identifying new information, evaluating its implications and determining its valuation impact is one of the most time-consuming and challenging aspects of real-world equity coverage. This is becoming ever more difficult and important as AI rapidly increases the quantity of new information to process. The strongest agent we evaluate on Frontier Financial Judgement matches all expert labels in only 52.4% of cases. We also find significant divergence in estimated false-positive rates among frontier agents, ranging from ~1% for GPT-5.6 Sol to ~32% for Claude Sonnet 4.6. To construct the benchmark and make it representative of real-world settings, we combine human-designed and labelled synthetic articles with live news articles and historical documents, creating 656 items for assessment. The resulting task requires agents to distinguish genuinely new, valuation-relevant financial information from stale, immaterial or misleading news under realistic conditions. We find substantial trade-offs among agent accuracy, cost, false positives and reliability that continue to hinder the reliable deployment of news-flow filtering in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。