arXiv:2509.02164cs.CV2025-09被引 5

首个跨帧全景问答数据集,提升全景场景理解能力

Omnidirectional Spatial Modeling from Correlated Panoramas

  • 构建跨帧相关全景图像的视觉问答数据集
  • 多模态大模型在问答任务中提升5.37%性能
  • 适合研究全景视觉与语言推理的学者

全方位场景理解对具身智能、自动驾驶和沉浸式环境等应用至关重要,但受限于360°图像中的几何畸变和复杂空间关系。现有方法仅关注单帧场景理解,忽略了跨帧相关全景图。为此,我们提出首个专注于整体360°场景中跨帧相关全景图视觉问答的基准数据集CFpano,包含超过2700张图像和8000多个问答对,涵盖选择题与开放题。基于此,我们进一步提出 extbf{methodname},一种通过组相对策略优化(GRPO)和定制奖励函数微调的多模态大语言模型,实现对跨帧相关全景图的鲁棒一致推理。在现有MLLM上进行基准测试,结果表明 extbf{methodname}在多项选择和开放问答任务中均达到领先性能,所有主要推理类别均优于强基线模型,总体提升+5.37%。分析验证了GRPO的有效性,并建立了新的全景场景理解基准。

原文摘要 · Abstract (English)

Omnidirectional scene understanding is vital for various downstream applications, such as embodied AI, autonomous driving, and immersive environments, yet remains challenging due to geometric distortion and complex spatial relations in 360° imagery. Existing omnidirectional methods achieve scene understanding within a single frame while neglecting cross-frame correlated panoramas. To bridge this gap, we introduce \textbf{CFpano}, the \textbf{first} benchmark dataset dedicated to cross-frame correlated panoramas visual question answering in the holistic 360° scenes. CFpano consists of over 2700 images together with over 8000 question-answer pairs, and the question types include both multiple choice and open-ended VQA. Building upon our CFpano, we further present \methodname, a multi-modal large language model (MLLM) fine-tuned with Group Relative Policy Optimization (GRPO) and a set of tailored reward functions for robust and consistent reasoning with cross-frame correlated panoramas. Benchmark experiments with existing MLLMs are conducted with our CFpano. The experimental results demonstrate that \methodname achieves state-of-the-art performance across both multiple-choice and open-ended VQA tasks, outperforming strong baselines on all major reasoning categories (\textbf{+5.37\%} in overall performance). Our analyses validate the effectiveness of GRPO and establish a new benchmark for panoramic scene understanding.

全景理解视觉问答多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。