arXiv:2608.11224cs.AIcond-mat.mtrl-sci2026-08

让AI记住材料研究经验,实现持续进化。

Harnessing agent memory to build lifelong AI partners for materials scientists

  • 构建可持久存储科学经验的自进化记忆框架
  • 在真实任务中使成功率接近翻倍,错误率下降92%
  • 适合需要长期积累与复用科研经验的研究者

材料研究依赖于积累的经验——有效的脚本、可信的流程、失败计算的警告以及将新问题与旧结果关联的判断。这些经验对可复现性和知识传递至关重要,却通常分散在笔记本、代码库、日志和个体记忆中,难以在AI代理间迁移。本文提出,应以持久记忆为核心,而非特定代理实现,设计终身材料科学AI伙伴。我们引入一个自进化记忆框架,将科学经验以可检查的事实和可执行技能形式存储,使观察结果、失败边界、流程和验证检查得以检索、修正并跨模型迁移。在三个计算场景中评估:在49个真实材料工具使用问题(共138个可执行子任务)中,记忆使GPT-5.2任务成功率几乎翻倍,无需更新模型参数;在元素固态物态方程计算中,记忆将波函数初始化失败转化为预执行防护机制,正确/部分正确/错误结果从22/1/4提升至25/2/0,避免92%重复错误;在13个实际材料模拟工作流中,记忆技能与失败事实使总轨迹负担(tokens)减半,第三轮后工具调用次数减少超一倍,同时保持能带隙、声子、空位和功函数分析的物理合理性。结果表明,代理记忆可作为持久的科学资产,是可移植、自我改进的材料研究经验记录,超越单一模型或代理架构的生命周期。

原文摘要 · Abstract (English)

Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed calculations or experiments, and judgement that links a new question to an old result. This experience is essential for reproducibility and knowledge transfer, yet it is usually fragmented across notebooks, repositories, job logs and individual memory, and it is rarely portable across artificial-intelligence agents. Here we argue that a lifelong AI partner for materials science can be designed around persistent memory rather than around a particular agent implementation. We introduce a self-evolving memory framework that stores scientific experience as inspectable facts and executable skills, so that observations, failure boundaries, protocols and validation checks can be retrieved, revised and migrated across models. We evaluate the idea in three computational settings that expose different layers of materials-research competence. In 49 real-world materials-tool-use questions comprising 138 executable subtasks, memory nearly doubles GPT-5.2 task success without model-parameter updates. In elemental-solid equation-of-state calculations, memory converts a wavefunction-initialization failure into a pre-execution guardrail, improving outcomes from 22/1/4 to 25/2/0 Correct/Partial/Error and avoiding 92% of repeated errors. In 13 practical material simulation workflows, remembered skills and failure facts halve the aggregate trace burden (tokens) and reduce tool calls by over a factor of two by the third round, while preserving physically meaningful outputs in band-gap, phonon, vacancy and work-function analyses. These results show that agent memory can serve as a durable scientific asset; a portable, self-improving record of materials-research experience that outlives any single model or agent stack.

AI助手材料科学记忆机制终身学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。