arXiv:2507.09155cs.CLcs.AI2025-07中稿 · Digital Discovery被引 1

评测大模型在晶体学问答中使用外部知识的能力,发现中小模型受益最大。

OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering

  • 构建217道专家精校的晶体学问题,分闭卷与开卷测试模型知识吸收能力。
  • 7B-70B参数模型从上下文获益最多,超大规模模型易出现性能饱和或干扰。
  • 专家审核的参考材料比AI生成效果好,内容质量远胜于文本长度。

我们提出OPENXRD,一个面向大语言模型(LLMs)和多模态大语言模型(MLLMs)在晶体学问答任务中的综合性评估框架。该框架聚焦于上下文吸收能力,即模型在推理时利用固定领域支持信息的能力。框架包含217道由专家精心设计的X射线衍射(XRD)问题,覆盖基础到高级晶体学概念,每道题均在闭卷(无上下文)与开卷(有上下文)条件下评估,后者引入由GPT-4.5生成并经晶体学专家优化的简洁参考段落。我们对74个最先进的LLM和MLLM(包括GPT-4、GPT-5、O系列、LLaVA、LLaMA、QWEN、Mistral、Gemini等系列)进行了基准测试,量化不同架构与规模下模型对外部知识的整合能力。结果显示,中等规模模型(7B–70B参数)从上下文获取的增益最大,而超大规模模型常出现饱和或干扰现象;在相同词元数量下,专家评审材料带来的提升显著高于AI生成材料,证实内容质量而非数量是性能关键。OPENXRD提供可复现的诊断基准,用于评估科学领域中的推理、知识融合与引导敏感性,并为未来多模态及检索增强型晶体学系统奠定基础。

原文摘要 · Abstract (English)

We introduce OPENXRD, a comprehensive benchmarking framework for evaluating large language models (LLMs) and multimodal LLMs (MLLMs) in crystallography question answering. The framework measures context assimilation, or how models use fixed, domain-specific supporting information during inference. The framework includes 217 expert-curated X-ray diffraction (XRD) questions covering fundamental to advanced crystallographic concepts, each evaluated under closed-book (without context) and open-book (with context) conditions, where the latter includes concise reference passages generated by GPT-4.5 and refined by crystallography experts. We benchmark 74 state-of-the-art LLMs and MLLMs, including GPT-4, GPT-5, O-series, LLaVA, LLaMA, QWEN, Mistral, and Gemini families, to quantify how different architectures and scales assimilate external knowledge. Results show that mid-sized models (7B--70B parameters) gain the most from contextual materials, while very large models often show saturation or interference and the largest relative gains appear in small and mid-sized models. Expert-reviewed materials provide significantly higher improvements than AI-generated ones even when token counts are matched, confirming that content quality, not quantity, drives performance. OPENXRD offers a reproducible diagnostic benchmark for assessing reasoning, knowledge integration, and guidance sensitivity in scientific domains, and provides a foundation for future multimodal and retrieval-augmented crystallography systems.

大模型评测晶体学知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。