arXiv:2505.14197cs.CV2025-05被引 15

首个全景视觉问答数据集与评测,提升模型对360°图像的理解能力。

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method

  • 基于GRPO设计三类奖励函数,优化全景图像推理过程。
  • 在360°场景中回答准确率提升6%,显著优于现有模型。
  • 适合研究沉浸式视觉理解、多模态大模型的开发者参考。

全景图像(ODIs)具有360°视场,为增强现实和具身AI等沉浸式应用提供了无与伦比的空间感知能力。然而,现有多模态大语言模型(MLLMs)对这类全景场景的理解与推理能力仍待探索。本文提出首个全景视觉问答数据集OmniVQA及基准评测,评估主流MLLMs在该任务上的表现,发现其在目标定位、特征提取和幻觉抑制方面存在显著局限,暴露出当前模型能力与全景理解需求之间的脱节,亟需针对360°图像的专用架构或训练方法。基于OmniVQA,我们进一步提出基于Qwen2.5-VL-Instruct的规则强化学习方法360-R1,通过改进组相对策略优化(GRPO),引入三类新奖励:推理过程相似性奖励、答案语义准确性奖励、结构化格式合规性奖励。在OmniVQA上的大量实验表明,所提方法在全景空间中实现+6%的性能提升。

原文摘要 · Abstract (English)

Omnidirectional images (ODIs), with their 360° field of view, provide unparalleled spatial awareness for immersive applications like augmented reality and embodied AI. However, the capability of existing multi-modal large language models (MLLMs) to comprehend and reason about such panoramic scenes remains underexplored. This paper addresses this gap by introducing OmniVQA, the first dataset and conducting the first benchmark for omnidirectional visual question answering. Our evaluation of state-of-the-art MLLMs reveals significant limitations in handling omnidirectional visual question answering, highlighting persistent challenges in object localization, feature extraction, and hallucination suppression within panoramic contexts. These results underscore the disconnect between current MLLM capabilities and the demands of omnidirectional visual understanding, which calls for dedicated architectural or training innovations tailored to 360° imagery. Building on the OmniVQA dataset and benchmark, we further introduce a rule-based reinforcement learning method, 360-R1, based on Qwen2.5-VL-Instruct. Concretely, we modify the group relative policy optimization (GRPO) by proposing three novel reward functions: (1) reasoning process similarity reward, (2) answer semantic accuracy reward, and (3) structured format compliance reward. Extensive experiments on our OmniVQA demonstrate the superiority of our proposed method in omnidirectional space (+6% improvement).

视觉问答全景图像强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。