arXiv:2604.08884cs.CVcs.AI2026-04

用三视图让大模型读懂高光谱图像,提升识别准确率。

Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Training-Free Reasoning Framework

论文配图:Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Training-Free Reasoning Framework
图 1 · 摘自论文原文
  • 将高光谱数据转为三视图输入,无需训练即可适配现有模型
  • 在17个模型上平均准确率从38.26%提升至40.35%
  • 适合关注遥感、材料分析等跨模态应用的研究者

多模态大语言模型(MLLM)在可见光图像理解中表现优异,但对可见光以外的光谱信息利用仍不充分。高光谱影像(HSI)包含丰富的材料与环境线索,但现有模型无法直接处理高维数据。为此,我们构建了HM-Bench基准,涵盖13类任务的19,337个问答对,来自2,178个高光谱样本。为使现有模型能使用HSI,我们提出VSR²框架——通过三个对齐视图表示:用于语义和空间上下文的RGB图像、基于PCA的主光谱变化图、以及定量光谱-空间证据结构化报告。在控制条件下对比仅用RGB与使用VSR²,17个代表性MLLM的平均准确率从38.26%提升至40.35%,不同模型和任务间差异显著。结果表明,高光谱信息对当前模型有用,但鲁棒的高光谱推理仍是开放挑战。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved strong performance on RGB image understanding, yet their ability to use spectral evidence beyond the visible range remains largely unexplored. Hyperspectral imagery (HSI) provides dense spectral measurements that reveal material and environmental cues unavailable in RGB, but current MLLMs cannot directly ingest high-dimensional HSI. To study this gap, we introduce HM-Bench, an evidence-grounded benchmark for hyperspectral image understanding with MLLMs. HM-Bench contains 19,337 question--answer pairs from 2,178 hyperspectral samples across 13 task categories, covering general perception, spectral reasoning, and spatial--spectral reasoning. To make HSI accessible to the native image--text interfaces of existing MLLMs, we further propose VSR$^{2}, a training-free Visual--Spectral--Report Reasoning framework. VSR^{2} represents each HSI sample with three aligned views: an RGB image for visual semantics and spatial context, a PCA-based image for dominant spectral variation, and a structured report for quantitative spectral--spatial evidence. Under a controlled RGB-only versus VSR^{2} evaluation protocol, experiments on 17 representative MLLMs show that HSI information improves average accuracy from 38.26% to 40.35%, with gains varying substantially across models and task categories. These results indicate that HSI information (beyond RGB) is useful for current MLLMs, while robust hyperspectral reasoning remains an open challenge.

高光谱多模态推理框架视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。