构建可追溯证据的智能助手,从多模态科学文献中自动提取结构化信息。
Building Agent Harnesses for Scientific Curation from Multimodal Sources

- 设计分阶段可审计流程,结合多模态工具与任务支架提升推理能力。
- 在黄金标准属性评分(GRAS)上达81.0,比前沿模型高23个百分点。
- 适合需要跨图文推理的科研自动化场景,尤其擅长高价值复杂属性提取。
科学发现工作流常依赖文献中的结构化整理,但现有智能体难以处理分散在长文本、密集表格和图表中的关键证据,且最终记录需跨多个证据片段进行推理而非简单复制。本文研究从多模态来源进行科学整理,提出Beaver——一种保留证据溯源的智能体协作框架。Beaver融合前沿智能体、多模态证据工具、任务支架与基于产物的自研研究机制,将整理过程变为分阶段、可审计的工作流,并支持迭代评估—诊断—修订循环,持久化的运行产物能暴露各阶段故障并指导框架优化。实验表明,Beaver在黄金标准属性评分(GRAS)上达到81.0,超越前沿智能体超过23个绝对百分点。消融实验显示,任务支架、多模态证据工具和溯源记录均对性能有显著贡献,属性层面分析表明,在需跨模态推理与归一化的高价值属性上收益最大。结果表明,对于包含多模态证据的科学文献整理,框架设计是决定智能体表现的核心因素。
原文摘要 · Abstract (English)
Scientific discovery workflows often depend on structured curation from the literature. This is difficult for current agents because the key evidence is scattered across long text, dense tables, and figures, and the final records often require reasoning across multiple evidence fragments rather than copying a single span. We study scientific curation from multimodal sources and introduce Beaver, an agent harness that extracts structured information from scientific papers while preserving provenance to the supporting evidence. Beaver combines a frontier agent with multimodal evidence tooling, task scaffolding, and artifact-grounded autoresearch. These components turn curation into a staged, auditable workflow and enable an iterative evaluate--diagnose--revise loop, where persistent run artifacts expose stage-localized failures and guide harness updates. Experiments show that Beaver reaches 81.0 on Gold-Referenced Attribute Score (GRAS), an attribute-level measure of agreement with gold curated records, outperforming frontier agents by over 23 absolute points. Ablations show that task scaffolding, multimodal evidence tooling, and provenance traces each contribute meaningfully to performance, while attribute-level analysis shows the largest gains on high-value attributes that require cross-modal reasoning and normalization. These results show that, for scientific curation from papers with multimodal evidence, harness design is a central determinant of agent performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。