arXiv:2607.19935cs.AI2026-07

用工具增强的AI系统,能精准发现并解释金属有机框架材料文件中的微小错误。

MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing

论文配图:MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing
图 1 · 摘自论文原文
  • 结合化学计算工具与强化学习,让AI基于真实数据生成可验证的错误解释。
  • 在四个基准上表现优于现有方法,错误检测率提升12%,解释可信度显著提高。
  • 适合材料科研人员和算法开发者,用于高精度材料数据质量审核。

大型金属有机框架(MOF)数据库通过晶格信息文件(CIF)支持模拟、筛选与机器学习。这些输入中细微的化学与结构错误会损害下游结果,且难以人工排查。大语言模型在计算化学中的进展为超越预测筛选、实现基于证据的细粒度诊断提供了可能。但仍有两大挑战:(i)细粒度归因能力有限:现有MOF专用验证器和机器学习模型仅提供固定检查、评分或粗略标签,缺乏基于证据的解释;(ii)CIF推理不可靠:直接使用大模型审计成本高且不准确,因化学证据隐含于原子坐标记录中,需进行几何、连接性、占据率、电荷等计算。两者均源于化学证据与语言模型解释间的耦合薄弱。本文提出MOF-Sleuth,一种基于强化学习的审计代理,包含确定性取证实验室与侦探推理引擎。实验室提取成分、几何、连接性、占据率、配位数及电荷等证据,侦探引擎利用这些证据生成有依据的解释、错误类型及二分类判断。奖励引导的强化学习将工具测量转化为解释层面的监督信号,不仅奖励最终答案,还奖励被引用的化学证据与证据支持的诊断。我们引入化学根基诊断(Chem-GD)评估指标,检验正确诊断是否由真实、相关的CIF衍生证据支撑。在四个基准上,MOF-Sleuth在基于大模型的方法和专门的机器学习方法中均达到领先性能,展示了在检测、归因与解释质量上的提升。

原文摘要 · Abstract (English)

Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information files (CIFs). Subtle chemical and structural errors in these inputs can compromise downstream results and hinder manual inspection. LLM advances in computational chemistry offer paths beyond predictive screening toward fine-grained diagnosis with evidence-grounded explanations. However, two challenges remain: (i) limited fine-grained attribution: MOF-specific validators and machine-learning models scale detection but provide fixed checks, readiness scores, or coarse labels rather than evidence-grounded explanations; and (ii) unreliable CIF reasoning: direct LLM auditing is costly and unreliable because chemical evidence is implicit across atom-site records and requires geometric, connectivity, occupancy, and charge calculations. Both stem from weak coupling between chemical evidence and language-model explanation. We introduce MOF-Sleuth, a reinforcement-guided CIF auditing agent with two modules: a deterministic Forensic Lab and a Sleuth reasoning engine. The Lab derives composition, geometry, connectivity, occupancy, coordination, and charge evidence, and Sleuth uses this evidence to produce an evidence-grounded explanation, error types, and a binary decision. Reward-guided reinforcement learning (RL) turns tool measurements into chemical explanation-level supervision, rewarding not only the final answer but also cited chemical evidence and evidence-supported diagnoses. We introduce Chemically Grounded Diagnosis (Chem-GD), a metric that assesses whether a correct diagnosis is explained by factual, relevant CIF-derived evidence. Across four benchmarks, MOF-Sleuth establishes state-of-the-art performance among LLM-based approaches and MOF-specific machine-learning methods, demonstrating gains in detection, attribution, and grounded explanation quality.

材料科学AI审计化学计算大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。