arXiv:2604.05557cs.CL2026-04被引 2

评测多轮多模态研究智能体的跨论文证据整合能力

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents

  • 设计多轮任务模拟真实科研流程,要求跨论文检索与整合图文证据
  • 顶尖模型在难题集上准确率仅29.23%,暴露现有系统短板
  • 适合评估可复现科研智能体,推动可信研究自动化

科学研究依赖多轮、多步骤的复杂工作流,需主动查阅文献、分析图表并整合多篇论文证据以对齐实验设置、支持可复现结论。现有基准未能系统评估主动检索、多证据融合及长期证据使用能力。本文提出EpiBench,一个基于片段式多轮多模态的研究工作流基准,模拟短期科研任务。给定研究任务后,智能体需在多轮中跨论文导航,对齐图表示例和表格数据,并利用记忆中的累积证据回答需跨论文比较与多图整合的客观问题。EpiBench引入过程级评估框架,实现对研究智能体的细粒度测试与诊断。实验显示,即使领先模型在难题集上准确率也仅为29.23%,表明多轮多证据研究工作流仍有巨大提升空间,为可验证、可复现的科研智能体提供评估平台。

原文摘要 · Abstract (English)

Scientific research follows multi-turn, multi-step workflows that require proactively searching the literature, consulting figures and tables, and integrating evidence across papers to align experimental settings and support reproducible conclusions. This joint capability is not systematically assessed in existing benchmarks, which largely under-evaluate proactive search, multi-evidence integration and sustained evidence use over time. In this work, we introduce EpiBench, an episodic multi-turn multimodal benchmark that instantiates short research workflows. Given a research task, agents must navigate across papers over multiple turns, align evidence from figures and tables, and use the accumulated evidence in the memory to answer objective questions that require cross paper comparisons and multi-figure integration. EpiBench introduces a process-level evaluation framework for fine-grained testing and diagnosis of research agents. Our experiments show that even the leading model achieves an accuracy of only 29.23% on the hard split, indicating substantial room for improvement in multi-turn, multi-evidence research workflows, providing an evaluation platform for verifiable and reproducible research agents.

多轮推理科研自动化多模态基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。