arXiv:2506.13051cs.CVcond-mat.mtrl-sci2025-06被引 5

用物理约束测试多模态模型在晶体推理中的泛化能力

Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning

  • 设计双基准测试:空间排除与成分排除,模拟不同泛化场景
  • 发现多数模型在成分外推时误差翻倍,且30%生成结构违反物理规律
  • 适合关注材料科学模型可靠性与物理一致性验证的研究者

评估晶体学推理的基座模型需要能分离泛化行为并强制执行物理约束的基准。本文引入一个多尺度多晶体数据集及两种基于物理的评估协议,用于压力测试多模态生成模型。空间排除基准从多样化数据集中隐藏特定半径内的超胞,实现对空间内插和外推的受控评估;成分排除基准剔除特定化学组成的样本,检验跨化学计量比的泛化能力。九个视觉-语言基座模型被输入晶体学图像与文本上下文,生成结构注释。评估指标包括:(i) 晶格参数与密度的相对误差,(ii) 物理一致性指数(惩罚体积违规),(iii) 幻觉分数(捕捉几何异常与无效空间群预测)。这些基准建立了可复现、基于物理的框架,用于评估大规模多模态模型的泛化性、一致性和可靠性。数据集与代码已开源:https://github.com/KurbanIntelligenceLab/StressTestingMMFMinCR。

原文摘要 · Abstract (English)

Evaluating foundation models for crystallographic reasoning requires benchmarks that isolate generalization behavior while enforcing physical constraints. This work introduces a multiscale multicrystal dataset with two physically grounded evaluation protocols to stress-test multimodal generative models. The Spatial-Exclusion benchmark withholds all supercells of a given radius from a diverse dataset, enabling controlled assessments of spatial interpolation and extrapolation. The Compositional-Exclusion benchmark omits all samples of a specific chemical composition, probing generalization across stoichiometries. Nine vision--language foundation models are prompted with crystallographic images and textual context to generate structural annotations. Responses are evaluated via (i) relative errors in lattice parameters and density, (ii) a physics-consistency index penalizing volumetric violations, and (iii) a hallucination score capturing geometric outliers and invalid space-group predictions. These benchmarks establish a reproducible, physically informed framework for assessing generalization, consistency, and reliability in large-scale multimodal models. Dataset and code are available at https://github.com/KurbanIntelligenceLab/StressTestingMMFMinCR.

晶体推理多模态模型物理约束基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。