arXiv:2607.02479cs.CV2026-07

解决全景图像搜索中视野受限问题,实现全局到局部的智能探索。

EAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$

论文配图:EAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$
图 1 · 摘自论文原文
  • 用全局先验引导搜索,逐步缩小局部范围,避免盲目查看。
  • 在14000+全景图上训练,目标检测准确率提升近8倍。
  • 适合做全景视觉导航、虚拟助手等需要空间推理的应用。

尽管多模态大语言模型在标准视觉理解任务中表现优异,但在360°全景环境中的主动视觉搜索仍面临根本性挑战。传统模型难以有效建模全景图特有的极坐标畸变和连续柱状拓扑结构,导致目标检测准确率下降。现有方法依赖碎片化局部视角,受限于固定初始位置且缺乏全局先验,易陷入短视、低效探索,且难以恢复丢失目标。为此,我们提出EAGLE-360框架,采用从全局到局部的主动探索策略:利用全局先验建立整体视角,迭代推理并逐步缩小搜索空间。通过引入坐标平移的位置编码机制(RoPE Rolling),无缝建模全景连续拓扑。构建了包含14,000+ 4K全景图与70,000+高质量视觉问答对话的大规模数据集。结合监督微调(SFT)与组相对策略优化(GRPO)的训练流程,有效激发复杂空间推理与工具调用能力。大量实验表明,EAGLE-360在360°视觉搜索任务上达到新基准,相比基础模型准确率提升近8倍,同时显著提高探索效率。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in standard visual understanding, adapting them for active visual search in 360$^\circ$ panoramic environments exposes fundamental limitations. Specifically, standard MLLMs struggle to effectively model inherent panoramic properties, such as severe polar distortion and continuous cylindrical topologies, which significantly degrades target detection accuracy. Consequently, existing panoramic search methods attempt to compensate by relying heavily on fragmented local viewpoints. Burdened by rigid initialization and a lack of global panoramic priors, these approaches suffer from myopic, inefficient exploration and struggle with robust error recovery when targets are out of view. To overcome these challenges, we propose EAGLE-360, a novel Embodied Active Global-to-Local Exploration framework. Rather than performing exhaustive local searches, EAGLE-360 leverages global priors to establish an initial holistic perspective, iteratively reasoning and progressively narrowing the search space. Architecturally, we adapt RoPE Rolling, a coordinate-shifting positional encoding mechanism, to seamlessly model the continuous topologies of panoramas. To facilitate this paradigm, we construct the large-scale EAGLE-360 dataset, comprising 14,000+ 4K panoramas and 70,000+ rounds of high-quality VQA dialogues. By employing a training pipeline that integrates Supervised Fine-Tuning (SFT) with Group Relative Policy Optimization (GRPO), we effectively elicit complex spatial reasoning and tool-calling capabilities. Extensive experiments demonstrate that EAGLE-360 establishes a new state-of-the-art for 360$^\circ$ visual search, achieving nearly an 8-fold increase in accuracy over the base model while significantly enhancing exploration efficiency.

全景搜索空间推理大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。