arXiv:2607.09195cs.AIcond-mat.mtrl-sci2026-07被引 1

让大模型科研过程可审计,通过显式假设演化提升可信度

Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents

论文配图:Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents
图 1 · 摘自论文原文
  • 设计可追踪的假设演化协议,将推理过程结构化
  • 在材料科学任务中实现完整假设-验证-修正循环
  • 适合需要可解释科研流程的研究者和审稿人

大型语言模型(LLM)代理正被寄予厚望,在人工智能驱动的科学发现中发挥核心作用。凭借广泛的知识、灵活的推理和工具使用能力,它们有望通过反复提出假设、测试并根据证据修正信念,自主探索与解决科学问题。然而,当前代理的假设、测试和信念更新隐藏于非结构化日志中,缺乏可审计机制。本文提出假设演化协议(Hypothesis Evolution Protocol, HEP),一种使假设生成、评估与演进成为显式、可审计操作的代理框架。在材料科学研究任务中,配备HEP的代理能够执行规划型代理所缺乏的假设-测试-证据-信念循环,跨研究问题具备泛化能力,且随着基础LLM能力提升,更充分地利用该协议。这些成果标志着向可审计的AI科学家迈进的重要一步,其科学推理过程可被检验、验证与延续。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scientific discovery. Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and solve scientific problems by repeatedly proposing hypotheses, testing them, and revising their beliefs in the light of the evidence. In current agents, however, these hypotheses, tests, and belief updates are buried in unstructured logs, and no mechanism lets the agent or the human researcher audit that process. Here we propose the Hypothesis Evolution Protocol (HEP), an agent harness that provides hypothesis generation, evaluation, and evolution as explicit, auditable operations. On materials-science research tasks, a HEP-equipped agent operates the hypothesis--test--evidence--belief cycle that planning-style agents lack, generalizes across research questions, and exploits the protocol more fully as the base LLM becomes more capable. These results mark a step toward auditable AI scientists, whose scientific reasoning can be inspected, verified, and built upon.

AI科研可解释性假设演化LLM代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。