arXiv:2605.29446cs.AI2026-05被引 1

评测视觉语言模型在晶体衍射峰索引中的表现,揭示当前技术仍远未解决该科学难题。

CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials

论文配图:CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials
图 1 · 摘自论文原文
  • 构建250样本晶体衍射基准数据集,结合图像与化学结构文本进行多步晶体学推理评测
  • 最佳模型仅达0.5888的雅各布逊分数和37.6%精确匹配率,多数模型低于0.50
  • 发现双峰情况易出错、过预测导致召回率虚高,文本信息无法弥补计算能力不足

从粉末XRD图谱中识别米勒指数需要现有跨模态基准尚未测试的能力:模型需从渲染的科学曲线中读取微窄峰位置,并将其与多步晶体学推理关联。我们提出CrystalXRD-Bench,一个包含250个样本的基准,源自10个公开晶体学数据库,任务为恢复某高强峰对应的所有完整HKL值。每个样本配对渲染的XRD图像、原始CIF文本及化学式,可并列分析视觉提取错误与推理错误。我们评估了七种视觉语言模型,最佳模型(GPT-5.4)的雅各布逊分数为0.5888,精确匹配率为37.6%,但六种模型仍低于0.50;该任务尚未解决。错误模式系统性显现:双峰情况尤其脆弱,高召回模型通过过度预测提升覆盖范围,而获取CIF文本未能缩小晶体学计算差距。除模型排名外,本基准还揭示了当前VLM在定量科学图像上的失效条件。所有数据与评估代码将公开。

原文摘要 · Abstract (English)

Miller-index identification from powder XRD patterns requires capabilities untested by existing multimodal benchmarks: the model must read a narrow peak location from a rendered scientific curve and then connect that observation to multi-step crystallographic reasoning. We introduce CrystalXRD-Bench, a 250-sample benchmark built from 10 public crystallographic databases for a single task: recover the full set of HKLs contributing to the highest-intensity peak in an XRD pattern. Each sample pairs the rendered XRD image with the source CIF text and chemical formula, so visual extraction errors and reasoning errors can be examined side by side. We evaluate seven vision-language models. The best Jaccard score is 0.5888 (GPT-5.4) with an exact-match rate of 37.6%, yet six of seven models remain below Jaccard 0.50; the task is far from solved. Error patterns vary systematically: double-peak cases are especially brittle, recall-heavy models gain coverage by over-predicting HKLs, and access to CIF text does not close the gap in crystallographic calculation. Alongside model rankings, the benchmark identifies the conditions under which current VLMs fail on quantitative scientific figures. All data and evaluation code will be publicly available.

XRD分析视觉语言模型晶体学科学图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。