arXiv:2608.17906cs.AIcs.MA2026-08

AutoResearch让科研自动执行更可靠,先有真见解再动手实验。

AutoResearch: Insight In, Hallucination Out

论文配图:AutoResearch: Insight In, Hallucination Out
图 1 · 摘自论文原文
  • 两阶段设计:先生成有依据的研究计划,再执行并验证。
  • 在RSICD上提升召回率至34.69,错误事件仅5次,远低于其他系统。
  • 适合需要高可信度自动化科研的学者与研发团队。

自主研究系统日益具备完成复杂研究流程的能力,但自动化并不保证科学性。我们提出AutoResearch,一种连接研究构思与执行的两阶段系统。在构思阶段,它融合新兴研究信号与领域知识,提取可迁移机制,通过多模型生成与交叉审查产出可验证的研究方案。在执行阶段,协同智能体将方案分解为实验,迭代实施与诊断,并在采纳结论前进行独立证据评估。在跨模态检索、系统优化和基准驱动机器学习等场景中,AutoResearch成功将想法转化为可衡量进展,识别并修正不可靠结果,基于证据决定是否继续、修改或终止研究方向。例如,在RSICD基准上,其生成方案将平均召回率从32.84提升至34.69,仅记录5个经审计确认的问题事件,显著优于其他系统的11–27次。这体现了‘先有洞察,后出结论’的可信研究流程:洞察入,幻觉出。

原文摘要 · Abstract (English)

Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.

自动化科研可信推理多智能体研究流程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。