评测材料大模型在不同尺度下的结构幻觉与推理能力,揭示精度不能代表真实可靠性。
SCALAR: Quantifying Structural Hallucination, Consistency, and Reasoning Gaps in Materials Foundation Models
- 构建跨尺度晶体结构基准,测试模型对纳米颗粒的推理解析能力。
- 发现显式推理可降低幻觉和误差,但常破坏结果一致性与有效性。
- 适合关注材料生成与可信推理的科研人员使用。
大型语言模型在材料科学推理中应用日益广泛,但其在物理结构分布变化下的行为仍不清晰。我们提出SCALAR(结构一致性与跨尺度逻辑),一个评估几何尺度泛化能力及其与结构幻觉、一致性和推理关系的基准。基于标准晶胞表示,模型需在从几个原子到超过18,000原子的尺度范围内,通过超胞扩展和几何截断生成的约10万种纳米颗粒结构上进行推理,所有结构均来自DFT验证的晶胞。SCALAR包含三项任务:(i) CIF到性质预测;(ii) 带显式物理依据的链式思考推理;(iii) 逆向检索:根据目标性质从候选结构中识别对应晶体。评估采用结构化指标,涵盖数值误差、幻觉、跨提示一致性、单调推理、输出合法性及检索遗憾。对多种基础模型的实验表明,显式推理虽常减少幻觉和误差,却频繁导致一致性或合法性失稳。结果表明,几何尺度泛化不可仅由准确率推断。
原文摘要 · Abstract (English)
Large language models are increasingly applied to materials science reasoning, yet their behavior under physically structured distribution shifts remains poorly understood. We introduce SCALAR (Structural Consistency And Logic Across Regimes), a benchmark for evaluating geometric scale generalization and its connection to structural hallucination, consistency, and reasoning in materials foundation models. Given canonical crystal representations, models must reason about derived nanoparticle structures obtained through supercell expansion and geometric truncation across length scales spanning a few atoms to over 18,000 atoms, totaling $\approx$100,000 structures from DFT-validated unit cells. SCALAR defines three tasks. (i) CIF to property prediction. (ii) A Chain-of-Thought variant with explicit physics-grounded reasoning. (iii) Inverse retrieval identifying crystals from candidates given target properties. Outputs are evaluated via structured metrics capturing numeric error, hallucination, cross-prompt consistency, monotonic reasoning, output validity, and retrieval regret. Experiments across diverse foundation models reveal large, model-dependent shifts under explicit reasoning, often reducing hallucination and error, but frequently destabilizing consistency or validity. These results demonstrate that geometric scale generalization cannot be inferred from accuracy alone. Supplementary materials are available at https://github.com/KurbanIntelligenceLab/SCALAR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。