构建全景推理新基准,推动多步全局空间推理能力发展
OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning

- 设计多步骤链式推理框架,连接全景证据与中间推理
- 涵盖14.3K训练数据和6.7K评估数据,支持全局空间一致性验证
- 适合研究全景视觉理解、具身智能与多模态模型的学者
多模态大语言模型在空间推理方面展现出潜力,但对全景图像这一新兴视觉模态的能力仍探索不足。全景图360°×180°的完整视域天然支持复杂的全局多步推理,这也是其在具身智能等应用中的核心优势。然而,现有全景基准大多聚焦依赖局部线索或单步/少步推理的简单任务,忽略了全景图的本质优势,未能充分挖掘其潜力。为此,我们提出OmniCoT,一个面向全景空间推理的评测套件,旨在使多模态大模型能利用全局证据并跨视角进行多步推理。该套件包含:OmniCoT-B(6.7K数据)用于评估,衡量答案准确率与推理质量;OmniCoT-Real(1K数据)为人工标注的真实世界子集,量化模拟到现实的差距;以及针对训练的OmniCoT-T(14.3K数据),配备结构化的逐步思维链标注,明确关联中间推理与全景证据。基于OmniCoT-T,我们引入OmniCoT-R1,采用两阶段训练策略以适应几何复杂的全景空间:监督微调(SFT)将推理锚定于全景证据(如方位角、距离),而GRPO惩罚几何不一致路径,强化360°空间一致性。通过OmniCoT,我们希望重新校准全景空间推理的难度,更契合全景图像的内在能力,推动该领域实质性进展。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the emerging visual modality of panoramic imagery. The full 360°$\times$180° field of view of panoramas essentially supports complex global multi-step reasoning, which is also the fundamental advantage of panoramas in applications such as embodied intelligence. However, existing panoramic benchmarks largely focus on simplistic queries that rely on local cues or single-/few-step reasoning, thereby ignoring the fundamental advantage of panoramas and failing to fully exploit their potential. To address this gap, we introduce OmniCoT, a panoramic spatial reasoning suite designed to enable MLLMs to use global evidence and perform multi-step inference across viewpoints. It includes OmniCoT-B (6.7K data) for evaluation, which measures both answer accuracy and reasoning quality, OmniCoT-Real (1K data) as a manually annotated real-world subset to quantify the Sim-to-Real gap. For training, OmniCoT-T (14.3K data) is purpose-built with structured stepwise Chain-of-Thought annotations that explicitly link intermediate reasoning steps to panoramic evidence. Based on OmniCoT-T, we introduce OmniCoT-R1 and adopt a two-stage training strategy tailored to the geometrically complex panoramic space, where Supervised Fine-tuning (SFT) anchors reasoning to panoramic evidence (e.g., bearings, proximity) and GRPO penalizes geometrically incoherent paths to consolidate global 360° spatial consistency. Through OmniCoT, we aim to recalibrate the difficulty of panoramic spatial reasoning to better align with the intrinsic capabilities of panoramic imagery, thereby fostering meaningful progress in this research area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。