用知识图谱模拟人类内省,提升科研智能体的推理与规划能力。
DAVIS: Planning Agent with Knowledge Graph-Powered Inner Monologue
- 构建带时序结构记忆的代理,实现基于模型的科学任务规划。
- 在9个基础科学主题中8项超越前人,在多跳问答任务上表现优异。
- 首次在RAG流程中引入交互式检索,适合需要深度推理的研究场景。
设计能在实验室环境中协助研究人员完成复杂科学任务的通用智能体,是当前人工智能研究的重要目标。与日常任务不同,科学任务对推理能力、环境时序理解及安全性要求更高,现有方法难以满足。为此,我们提出DAVIS。不同于传统检索增强生成(RAG)方法,DAVIS引入结构化与时间记忆机制,支持基于模型的规划;同时采用类人类内省的多轮交互式检索系统,增强对过往经验的推理能力。DAVIS在ScienceWorld基准测试中,于9个基础科学主题中的8个上显著优于以往方法;其世界模型在著名的HotpotQA和MusiqueQA多跳问答数据集上也表现出色。据我们所知,DAVIS是首个在RAG流程中采用交互式检索的智能体。
原文摘要 · Abstract (English)
Designing a generalist scientific agent capable of performing tasks in laboratory settings to assist researchers has become a key goal in recent Artificial Intelligence (AI) research. Unlike everyday tasks, scientific tasks are inherently more delicate and complex, requiring agents to possess a higher level of reasoning ability, structured and temporal understanding of their environment, and a strong emphasis on safety. Existing approaches often fail to address these multifaceted requirements. To tackle these challenges, we present DAVIS. Unlike traditional retrieval-augmented generation (RAG) approaches, DAVIS incorporates structured and temporal memory, which enables model-based planning. Additionally, DAVIS implements an agentic, multi-turn retrieval system, similar to a human's inner monologue, allowing for a greater degree of reasoning over past experiences. DAVIS demonstrates substantially improved performance on the ScienceWorld benchmark comparing to previous approaches on 8 out of 9 elementary science subjects. In addition, DAVIS's World Model demonstrates competitive performance on the famous HotpotQA and MusiqueQA dataset for multi-hop question answering. To the best of our knowledge, DAVIS is the first RAG agent to employ an interactive retrieval method in a RAG pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。