用持续演化的知识图谱实现智能科研全流程监督。
AI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World Model
- 构建动态研究世界模型,统一多智能体认知基础。
- 自动发现方法缺陷并验证其在不同基准上的表现差异。
- 跨领域搜索解决方案,支持自我修正与迭代优化。
现有自动化研究系统以无状态、线性流程运行——生成结果时不保留对研究环境的持久理解。它们顺序处理论文,提出想法缺乏结构化差距分析,且缺乏智能体间验证、质疑或相互修正的机制。我们提出 extbf{AI-Supervisor},一个由专用智能体组成的多智能体协同框架,通过自主探索和对研究知识的自我修正更新,实现从文献综述到差距发现、方法开发、评估与论文撰写等环节的全流程人工智能研究监督。不同于线性流水线,AI-Supervisor 维护一个持续演进的 extit{研究世界模型},采用知识图谱实现,捕捉方法、基准、已知局限与未探索空白,作为所有智能体共享记忆,使智能体能基于结构化理解开展探索与深化。框架引入三项架构贡献:(1) 结构化差距发现,将方法分解为核心模块,验证其在多个基准上的性能,并映射各模块产生的具体缺口;(2) 自我修正发现循环,探究模块在特定问题上成功或失败的原因,识别基准是否隐含偏差,评估协议是否仍适用于新兴挑战;(3) 自我改进开发循环,通过跨领域机制搜索,迭代定位失效模块并从其他科学领域中寻找解决方案。所有智能体均遵循共识机制,独立发现需经验证后方可写入研究世界模型。
原文摘要 · Abstract (English)
Existing automated research systems operate as stateless, linear pipelines -- generating outputs without maintaining any persistent understanding of the research landscape they navigate. They process papers sequentially, propose ideas without structured gap analysis, and lack mechanisms for agents to verify, challenge, or refine each other's findings. We present \textbf{AI-Supervisor}, a multi-agent orchestration framework where specialized agents provide end-to-end AI research supervision driven by human interests -- from literature review through gap discovery, method development, evaluation, and paper writing -- through autonomous exploration and self-correcting updates of research knowledge. Unlike sequential pipelines, AI-Supervisor maintains a continuously evolving \emph{Research World Model}, implemented as a Knowledge Graph, that captures methods, benchmarks, known limitations, and unexplored gaps, serving as shared memory across all agents and enabling agents to explore and build upon a structured understanding of the research landscape. The framework introduces three architectural contributions: (1) \emph{structured gap discovery} that decomposes methods into core modules, validates their performance across benchmarks, and maps the specific gaps each module creates; (2) \emph{self-correcting discovery loops} that probe why modules succeed on certain problems and fail on others, whether benchmarks carry hidden biases, and whether evaluation protocols remain adequate for emerging challenges; and (3) \emph{self-improving development loops} governed by cross-domain mechanism search that iteratively targets failing modules by finding solutions from other scientific fields. All agents operate under a \emph{consensus mechanism} where independent findings are corroborated before being committed to the Research World Model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。