arXiv:2603.09573cs.CV2026-03中稿 · CVPR被引 5

用全景图提升视觉语言模型的全局理解能力,应对复杂场景挑战

More than the Sum: Panorama-Language Models for Adverse Omni-Scenes

  • 提出全景-语言建模新范式,统一处理360度图像
  • 在包含遮挡和事故的全景问答数据集上表现更优
  • 可插拔模块让旧模型无需重训即可处理全景图

现有视觉语言模型针对针孔成像设计,通过拼接多窄视图来构建全景理解,但忽略了单张全景图固有的整体空间与上下文关系。本文提出全景-语言建模(PLM)范式,实现超越其窄视图组件之和的360°视觉语言推理。同时发布大规模全景视觉问答数据集PanoVQA,涵盖遮挡与交通事故等恶劣全景场景,支持全面推理。为建立PLM基础,开发可插拔的全景稀疏注意力模块,使现有针孔模型无需重新训练即可处理等距柱状投影全景图。大量实验表明,该方法在复杂全景场景下具备更强鲁棒性与整体推理能力,实现超越各部分之和的理解。

原文摘要 · Abstract (English)

Existing vision-language models (VLMs) are tailored for pinhole imagery, stitching multiple narrow field-of-view inputs to piece together a complete omni-scene understanding. Yet, such multi-view perception overlooks the holistic spatial and contextual relationships that a single panorama inherently preserves. In this work, we introduce the Panorama-Language Modeling (PLM)paradigm, a unified $360^\circ$ vision-language reasoning that is more than the sum of its pinhole counterparts. Besides, we present PanoVQA, a large-scale panoramic VQA dataset that involves adverse omni-scenes, enabling comprehensive reasoning under object occlusions and driving accidents. To establish a foundation for PLM, we develop a plug-and-play panoramic sparse attention module that allows existing pinhole-based VLMs to process equirectangular panoramas without retraining. Extensive experiments demonstrate that our PLM achieves superior robustness and holistic reasoning under challenging omni-scenes, yielding understanding greater than the sum of its narrow parts. Project page: https://github.com/InSAI-Lab/PanoVQA.

全景视觉视觉语言模型稀疏注意力多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。