评测自动研究的科研过程是否像真人,发现当前模型仅达67.9分。
ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

- 构建模仿人类研究者的行为评估框架,从答案匹配转向过程还原。
- 11个顶尖框架最高仅得67.9分,显示与真实科研流程存在显著差距。
- 可诊断研究全流程,适合训练下一代自主科研系统。
Auto-Research 的快速发展暴露了一个根本性评估难题:如何衡量其研究轨迹在对齐性、逻辑连贯性和演化完整性方面与人类研究行为的一致性?我们提出 ARAC-Bench:一种模仿研究者的评估框架,目标从匹配最终答案转变为复现高质量的人类研究过程。该框架包含两个协同组件:学术认知能力系统,首次将隐式审稿人经验转化为阶段校准、可量化的评分标准;以及三阶段能力诊断协议,将研究过程在严格模块化约束下分解为三个可追踪、相互独立的维度:提案(Proposal)、实验(Experiment)和综合(Synthesis)。对11个最先进框架的系统评估显示,最佳对齐得分为67.9/100,揭示了模拟严谨人类方法学的巨大差距。与博士候选人排名的验证显示相关系数达0.8141,证实ARAC-Bench能可靠反映研究者真正重视的维度。ARAC-Bench不仅提供细粒度诊断工具,还可作为训练下一代自主研究系统的可扩展奖励信号。
原文摘要 · Abstract (English)
The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。