构建科学光谱理解基准,评估大模型对复杂谱图的问答能力。
SpecVQA: A Benchmark for Spectral Understanding and Visual Question Answering in Scientific Images

- 设计基于7类典型谱图的问答数据集,支持科学信息提取与推理。
- 包含620张图、3100个问答对,源自同行评审文献。
- 提出采样与插值重建方法,有效降低令牌长度并提升性能。
光谱是常见但信息密集的科学图像形式,因其非结构化和领域专属性,给多模态大语言模型(MLLMs)带来巨大挑战。本文提出SpecVQA,一个专业的科学图像基准,用于评估多模态模型在科学光谱理解方面的能力,涵盖7种代表性光谱类型,并配有专家标注的问答对。该基准旨在实现两大目标:科学光谱问答评估及对应底层任务评估。SpecVQA共包含620幅图像和3100个问答对,数据源自同行评审文献,覆盖直接信息抽取与领域特定推理。为在保留曲线关键特征的同时有效减少令牌长度,我们提出一种光谱数据采样与插值重建方法。消融实验表明,该方法在所提基准上显著提升模型性能。我们在该基准上测试了主流MLLMs在科学光谱理解方面的能力,并发布了排行榜。本工作是提升多模态大模型光谱理解能力的重要一步,也为扩展视觉-语言模型至更广泛的科学研究与数据分析提供了可行方向。
原文摘要 · Abstract (English)
Spectra are a prevalent yet highly information-dense form of scientific imagery, presenting substantial challenges to multimodal large language models (MLLMs) due to their unstructured and domain-specific characteristics. Here we introduce SpecVQA, a professional scientific-image benchmark for evaluating multimodal models on scientific spectral understanding, covering 7 representative spectrum types with expert-annotated question-answer pairs. The aim comprises two aspects: spectra scientific QA evaluation and corresponding underlying task evaluation. SpecVQA contains 620 figures and 3100 QA pairs curated from peer-reviewed literature, targeting both direct information extraction and domain-specific reasoning. To effectively reduce token length while preserving essential curve characteristics, we propose a spectral data sampling and interpolation reconstruction approach. Ablation studies further confirm that the approach achieves substantial performance improvements on the proposed benchmark. We test the capability of prominent MLLMs in scientific spectral understanding on our benchmark and present a leaderboard. This work represents an essential step toward enhancing spectral understanding in multimodal large models and suggests promising directions for extending visual-language models to broader scientific research and data analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。