arXiv:2603.16179cs.CVcs.AI2026-03被引 3

提出首个360°图像理解基准与免训练方法,解决全景图感知难题。

360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method

  • 构建7K分辨率全景图基准,涵盖七项任务
  • 免训练框架提升基线模型表现,支持高分辨率推理
  • 适合虚拟现实、智能导航等场景应用

多模态大语言模型(MLLMs)在常规图像理解与推理方面表现优异,但对360°图像的感知仍缺乏深入研究。与常规图像不同,360°图像捕捉完整环境,支持整体空间推理,但也带来几何失真和复杂空间关系等挑战。为此,我们引入360Bench,一个基于7K分辨率360°图像的视觉问答(VQA)基准,包含七个代表性(子)任务,标注由人工精心完成。利用360Bench,我们系统评估了七种MLLMs及六种增强方法,揭示其在360°图像感知中的不足。为应对这些挑战,我们提出Free360——一种免训练的基于场景图的框架,用于高分辨率360°图像问答。该方法将推理过程分解为模块化步骤,针对每一步应用自适应球面图像变换,并将结果无缝整合到统一图表示中以生成答案。实验表明,Free360持续提升基线模型性能,为360°图像问答提供强大免训练解决方案。代码与数据集将在论文录用后公开。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown impressive abilities in understanding and reasoning over conventional images. However, their perception of 360° images remains largely underexplored. Unlike conventional images, 360° images capture the entire surrounding environment, enabling holistic spatial reasoning but introducing challenges such as geometric distortion and complex spatial relations. To comprehensively assess MLLMs' capabilities to perceive 360° images, we introduce 360Bench, a Visual Question Answering (VQA) benchmark featuring 7K-resolution 360° images, seven representative (sub)tasks with annotations carefully curated by human annotators. Using 360Bench, we systematically evaluate seven MLLMs and six enhancement methods, revealing their shortcomings in 360° image perception. To address these challenges, we propose Free360, a training-free scene-graph-based framework for high-resolution 360° VQA. Free360 decomposes the reasoning process into modular steps, applies adaptive spherical image transformations to 360° images tailored to each step, and seamlessly integrates the resulting information into a unified graph representation for answer generation. Experiments show that Free360 consistently improves its base MLLM and provides a strong training-free solution for 360° VQA tasks. The source code and dataset will be publicly released upon acceptance.

360图像多模态免训练视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。