arXiv:2602.22971cs.AI2026-02

构建用于扫描探针显微镜的高精度自动化评测基准,突破科研领域AI评估瓶颈。

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy

  • 全自动合成数据,结合锚点筛除技术高效提取图文对。
  • 提出严格不完美惩罚F1评分,量化模型在复杂物理场景中的推理边界。
  • 可揭示模型性格特征,适合评估科研场景下大模型的真实能力。

随着大语言模型在通用推理任务中取得突破,其在专业科学领域的表现仍因数据污染、复杂度不足和人力成本高昂而存在显著差距。本文提出SPM-Bench,一个专为扫描探针显微镜(SPM)设计的博士级多模态基准。我们开发了全自动化数据合成流程,确保数据权威性与低成本并存。通过采用锚点门控筛除(AGS)技术,从2023至2025年发表的arXiv及期刊论文中高效提取高质量图像-文本对。利用混合云-本地架构,由视觉语言模型仅返回空间坐标“llbox”以实现局部高保真裁剪,极大节省令牌消耗同时保持数据纯净度。为客观评估模型性能,引入严格不完美惩罚F1(SIP-F1)评分,不仅建立严谨的能力层级,首次量化模型“性格”(保守型、激进型、赌徒型或明智型)。通过关联模型置信度与感知难度,揭示当前AI在复杂物理场景下的真实推理边界。这些发现使SPM-Bench成为可推广的自动化科学数据合成范式。

原文摘要 · Abstract (English)

As LLMs achieved breakthroughs in general reasoning, their proficiency in specialized scientific domains reveals pronounced gaps in existing benchmarks due to data contamination, insufficient complexity, and prohibitive human labor costs. Here we present SPM-Bench, an original, PhD-level multimodal benchmark specifically designed for scanning probe microscopy (SPM). We propose a fully automated data synthesis pipeline that ensures both high authority and low-cost. By employing Anchor-Gated Sieve (AGS) technology, we efficiently extract high-value image-text pairs from arXiv and journal papers published between 2023 and 2025. Through a hybrid cloud-local architecture where VLMs return only spatial coordinates "llbox" for local high-fidelity cropping, our pipeline achieves extreme token savings while maintaining high dataset purity. To accurately and objectively evaluate the performance of the LLMs, we introduce the Strict Imperfection Penalty F1 (SIP-F1) score. This metric not only establishes a rigorous capability hierarchy but also, for the first time, quantifies model "personalities" (Conservative, Aggressive, Gambler, or Wise). By correlating these results with model-reported confidence and perceived difficulty, we expose the true reasoning boundaries of current AI in complex physical scenarios. These insights establish SPM-Bench as a generalizable paradigm for automated scientific data synthesis.

大模型评测科学计算多模态数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。