arXiv:2607.02497cs.CV2026-07

让智能体主动转头找物体,实现全景指代分割。

Seek to Segment: Active Perception for Panoramic Referring Segmentation

论文配图:Seek to Segment: Active Perception for Panoramic Referring Segmentation
图 1 · 摘自论文原文
  • 用视觉-语言模型+空间记忆构建全景感知能力
  • 搜索效率提升40%,分割准确率超现有方法
  • 适合需要自主探索的机器人、虚拟助手场景

现有指代分割模型被动处理固定视角的静态图像,难以应用于具身智能中连续360°环境下的主动感知。为此,我们提出新任务:主动全景指代分割(APRS)。在此设定下,智能体需调整视角(Δθ, Δϕ)在360°环境中主动探索,定位用户指令指定的物体并进行分割。为此,我们提出PanoSeeker——一种基于记忆增强的智能体,通过将序列局部观测逐步融合为统一的360°表示,实现高效非冗余搜索。该模型结合视觉-语言模型与显式空间视觉记忆EgoSphere,支持规划最优探索路径。找到目标后,智能体执行主动视角对齐并输出分割掩码。我们还构建了专家标注的搜索轨迹数据集,包含记忆时间线,用于监督微调和强化学习后训练以优化探索效率。在新建立的APRS基准上的大量实验表明,PanoSeeker在搜索效率和分割准确率上均显著优于适配的最先进基线方法。

原文摘要 · Abstract (English)

Existing referring segmentation models passively process static images captured from fixed perspectives, limiting their applicability in Embodied AI, where agents must perform active perception in the continuous 360$^\circ$ environments. To bridge this gap, we introduce a novel task: Active Panoramic Referring Segmentation (APRS). In this setting, an agent is required to adjust its viewing direction ($Δθ, Δϕ$) to explore the 360$^\circ$ environment, seeking the object specified by a user instruction for segmentation. To tackle this challenging task, we propose PanoSeeker, a memory-augmented agent for efficient APRS. Rather than relying on heuristic scanning, PanoSeeker integrates a Vision-Language Model (VLM) with EgoSphere, an explicit spatial visual memory. By progressively integrating sequential local observations into a unified 360$^\circ$ representation, EgoSphere enables the agent to plan efficient and non-redundant search trajectories. Once the target is found, the agent performs active viewpoint alignment and outputs the segmentation mask. Furthermore, we curate an expert-annotated search trajectory dataset with memory timelines for Supervised Fine-Tuning, followed by Reinforcement Learning post-training to explicitly optimize PanoSeeker's exploration efficiency. Extensive experiments on our newly established APRS benchmark demonstrate that PanoSeeker achieves superior search efficiency and segmentation accuracy, significantly outperforming adapted state-of-the-art baselines.

具身智能主动感知全景分割视觉记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。