arXiv:2606.00644cs.AI2026-06

测试大模型能否基于历史数据预判AI研究未来方向

ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment

论文配图:ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
图 1 · 摘自论文原文
  • 构建时间控制的基准,模拟未来决策场景
  • 多领域500个任务验证,发现证据与判断常脱节
  • 适合评估科研智能体的决策能力,尤其关注逻辑一致性

AI研究常需在缺乏未来证据时做出决策:攻击哪个瓶颈、推进哪个方向、项目如何定位。我们提出ForeSci,一个时间可控的基准,用于评估大模型智能体能否基于历史证据进行前瞻性研究判断。ForeSci包含四个快速发展的AI领域中的500个任务,涵盖四类决策类型。每个任务配有截止时间对齐的离线知识库;生成阶段隐藏截止后论文,仅用于验证。为避免随机预测,任务源自截止前的分类分支和证据信号,且答案生成模型均早于任务截止时间。我们在四种基础模型上评估原生LLM、混合RAG及三种科研智能体适配方案。结果表明,显式组织证据可提升可追溯性和事实支持,但增益高度依赖决策类型。诊断发现存在普遍的证据-决策脱节现象:智能体可能引用相关证据,却预测错误的研究对象。ForeSci将前瞻性AI研究判断转化为可控制的基准,用于评估科研智能体作为决策系统的性能。

原文摘要 · Abstract (English)

AI research often requires decisions before future evidence exists: which bottleneck to attack, which direction to pursue, or where a project should be positioned. We introduce ForeSci, a temporally controlled benchmark for evaluating whether LLM agents can make such forward-looking research judgements from historical evidence. ForeSci contains 500 tasks across four fast-moving AI domains and four decision families. Each task is paired with a cutoff-aligned offline knowledge base; post-cutoff papers are hidden during generation and used only for validation. To avoid random future-event prediction, tasks are derived from pre-cutoff taxonomy branches and evidence signals, and answer-generation backbones are selected to precede the task cutoffs. We evaluate native LLMs, Hybrid RAG, and three research-agent adaptations across four backbones. Results show that explicit evidence organization improves traceability and factual support, but gains depend strongly on the decision family. Diagnostics reveal a recurring evidence-decision decoupling: agents may cite relevant evidence while forecasting the wrong research object. ForeSci turns forward-looking AI research judgement into a controlled benchmark for evaluating research agents as decision-making systems.

智能体评估前瞻决策科研辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。