arXiv:2605.10576cs.CVcs.AI2026-05

首个面向遥感图像低级视觉感知的诊断基准,让大模型学会识别并描述遥感图像失真。

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models

论文配图:SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models
图 1 · 摘自论文原文
  • 构建基于物理的分级分类体系,覆盖22类遥感退化类型
  • 测试29个主流大模型,发现其对遥感退化存在误判与描述偏差
  • 提供可解释的诊断数据,适合遥感与多模态模型研究者使用

低级视觉感知是可靠遥感(RS)图像分析的基础,但现有图像质量评估(IQA)方法输出不可解释的标量分数,无法刻画由物理机制驱动的遥感退化,与遥感专家的实际诊断需求严重脱节。尽管视觉-语言模型(VLMs)可通过语言描述实现可解释的IQA,但其视觉先验严重偏向地面自然图像。因此,现有方法尚不足以验证VLMs能否克服领域差距以感知并描述遥感退化。为此,我们提出 extbf{SenseBench},首个专用于遥感低级视觉感知与描述的诊断基准。基于物理驱动的分层分类体系,统一非参考与参考两种范式,包含超过10,000个精心标注的样本,覆盖6大类、22细粒度的遥感退化类别。设计两种互补评估协议:客观低级视觉 extit{感知} 与主观诊断 extit{描述}。对29个前沿VLMs的全面评估揭示了领域先验偏移、多退化混淆、 extit{流畅性错觉} 及 extit{感知-描述倒置} 等现象。我们期望 SenseBench 能为遥感低级感知中VLMs的发展提供可靠的评测平台与高质量诊断数据。代码与数据集见:https://github.com/Zhong-Chenchen/SenseBench。

原文摘要 · Abstract (English)

Low-level visual perception underpins reliable remote sensing (RS) image analysis, yet current image quality assessment (IQA) methods output uninterpretable scalar scores rather than characterizing physics-driven RS degradations, deviating markedly from the diagnostic needs of RS experts. While Vision-Language Models (VLMs) present a compelling alternative by delivering language-grounded IQA, their visual priors are heavily biased toward ground-level natural images. Consequently, whether VLMs can overcome this domain gap to perceive and articulate RS artifacts remains insufficiently studied. To bridge this gap, we propose \textbf{SenseBench}, the first dedicated diagnostic benchmark for RS low-level visual perception and description. Driven by a physics-based hierarchical taxonomy that unifies both non-reference and reference-based paradigms, SenseBench features over 10K meticulously curated instances across 6 major and 22 fine-grained RS degradation categories. Specifically, two complementary protocols are designed for evaluation: objective low-level visual \textit{perception} and subjective diagnostic \textit{description}. Comprehensive evaluation of 29 state-of-the-art VLMs reveals not only skewed domain priors and multi-distortion collapse, but also \textit{fluency illusion} and a \textit{perception-description inversion} effect. We hope SenseBench provides a robust evaluation testbed and high-quality diagnostic data to advance the development of VLMs in RS low-level perception. Code and datasets are available \href{https://github.com/Zhong-Chenchen/SenseBench}{\textcolor{blue}{here}}.

遥感视觉语言模型图像质量评估诊断基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。