arXiv:2605.18791eess.IVcs.CV2026-05

构建大规模多模态光谱基准SpecX,统一评估专业模型与多模态大模型性能

SpecX: A Large-Scale Benchmark for Multi-Modal Spectroscopy and Cross-Paradigm Evaluation

论文配图:SpecX: A Large-Scale Benchmark for Multi-Modal Spectroscopy and Cross-Paradigm Evaluation
图 1 · 摘自论文原文
  • 构建包含170万分子的多模态光谱数据集,覆盖核磁、红外等七种谱型
  • 专业模型擅长信号级建模,多模态大模型长于高层推理但缺乏谱图精确对齐
  • 适合光谱智能、药物发现和基础模型研究者使用

现有光谱基准在规模、模态对齐和评估范围上存在局限,通常仅针对专用模型或多模态语言模型(MLLMs)。我们提出SpecX,一个支持跨范式评估的大规模多模态光谱基准。SpecX包含170万分子,涵盖1H-NMR、13C-NMR、HSQC、IR、MS、UV、Raman和FL等多种光谱模态,并分为三个层级:用于预训练的大规模数据集、用于基准测试的对齐多谱子集,以及用于评估的高质量实验子集。该基准支持分子解析、谱图模拟和谱图理解等任务,可统一评估专用光谱模型与MLLMs。实验表明,专用模型在信号级建模上表现优异,而MLLMs在高层推理上具有优势,但缺乏精确的谱图语义对齐。SpecX为光谱智能建立了统一评估标准,凸显了开发谱图原生基础模型的必要性。

原文摘要 · Abstract (English)

Existing spectral benchmarks are limited in scale, modality alignment, and evaluation scope, and typically focus on either specialized models or multimodal language models (MLLMs). We introduce SpecX, a large-scale benchmark for multi-modal spectroscopy with cross-paradigm evaluation. SpecX contains 1.7M molecules with diverse spectral modalities, including NMR (1H, 13C, HSQC), IR, MS,UV,Raman and FL, and is organized into three tiers: a large-scale dataset for pretraining, an aligned multi-spectral subset for benchmarking, and a high-quality experimental subset for evaluation. SpecX supports a range of tasks such as molecular elucidation, spectrum simulation, and spectral understanding, and enables unified evaluation across both specialized spectral models and MLLMs. Experiments show that specialized models excel at signal-level modeling, while MLLMs exhibit strengths in high-level reasoning but lack precise spectral grounding. SpecX establishes a unified benchmark for spectral intelligence and highlights the need for spectrum-native foundation models.

光谱分析多模态基准测试基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。