AI能理解科学原理,却难预测未来突破。
Scientific reasoning does not reliably translate into scientific forecasting in frontier AI

- 构建跨八学科的时序科学预测评估集CUSP,测试模型前瞻性能力。
- 六款前沿AI模型虽能识别合理机制,但预测准确率接近随机水平。
- 模型常高估成果发布时间,且方案与实际进展偏差大,适合科研决策者关注。
AI系统越来越多地用于支持前瞻性的科学判断,但其能否对未来的科学进展形成可靠预期仍不明确。本文通过引入跨八大学科、时间锚定的事件级科学预测评估套件CUSP,研究此问题。在六款前沿AI模型中,观察到显著的预测性能不对称性及系统性误差模式:模型常能识别未来科学突破的合理机制,但在可行性评估上表现接近随机,生成的解决方案策略与实际实现进展关联性弱,且系统性地将科学突破预测为晚于公开可观察时间。提供额外的截断前科学知识可提升性能,但无法消除这些预测局限。结果表明,当前AI具备较强的回溯性科学能力,但前瞻预测能力有限。因此,在科研优先级设定与科学决策中部署AI时,应将科学预测能力作为独立维度进行评估。
原文摘要 · Abstract (English)
AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expectations about future scientific advances. Here we show that strong scientific reasoning does not reliably translate into accurate forecasting of future scientific advances. To study this question, we introduce CUSP, a temporally grounded evaluation suite for event-level scientific forecasting across eight scientific disciplines. Across six frontier AI models, we observe a striking asymmetry in forecasting performance together with systematic error patterns. Models often identify plausible mechanisms underlying future scientific advances, yet perform near chance on feasibility assessment, generate solution strategies that only weakly align with realized advances, and systematically predict scientific advances later than they become publicly observable. Providing additional pre-cutoff scientific knowledge improves performance but does not eliminate these forecasting limitations. These findings suggest that current AI systems possess substantial retrospective scientific competence but limited forward-looking predictive capability. Scientific forecasting should therefore be evaluated as a complementary dimension of AI scientific capability when deploying AI systems for research prioritization and scientific decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。