arXiv:2511.20351cs.CV2025-11被引 16

让机器人像人一样转动头颅360度搜寻目标,挑战真实场景下的视觉搜索难题。

Thinking in 360°: Humanoid Visual Search in the Wild

  • 构建可转动头部的类人智能体,在全景图像中主动搜索物体与路径。
  • 开源模型经优化后成功率提升三倍,但路径搜索仍不足25%。
  • 适用于需要空间推理的真实复杂场景,如交通枢纽和大型商场。

人类通过头部与眼球协同运动,在360°环境中高效搜索视觉信息。然而,现有视觉搜索方法多局限于静态图像,忽视了物理实体与三维世界的交互。为此,我们提出类人视觉搜索,让机器人在沉浸式360°全景世界中主动旋转头部进行搜索。为研究复杂现实场景下的视觉搜索,我们构建了H* Bench基准,涵盖交通枢纽、大型零售空间、城市街道等真实场景,需具备高级视觉-空间推理能力。实验表明,即使顶级闭源模型在物体与路径搜索中成功率也仅约30%。通过后训练技术,我们将开源模型Qwen2.5-VL的成功率提升三倍以上:物体搜索从14.83%升至47.38%,路径搜索从6.44%升至24.94%。路径搜索天花板较低,揭示其对空间常识的高要求。结果表明,尽管进展显著,但要实现能无缝融入日常生活的多模态大模型智能体,仍面临巨大挑战。

原文摘要 · Abstract (English)

Humans rely on the synergistic control of head (cephalomotor) and eye (oculomotor) to efficiently search for visual information in 360°. However, prior approaches to visual search are limited to a static image, neglecting the physical embodiment and its interaction with the 3D world. How can we develop embodied visual search agents as efficient as humans while bypassing the constraints imposed by real-world hardware? To this end, we propose humanoid visual search where a humanoid agent actively rotates its head to search for objects or paths in an immersive world represented by a 360° panoramic image. To study visual search in visually-crowded real-world scenarios, we build H* Bench, a new benchmark that moves beyond household scenes to challenging in-the-wild scenes that necessitate advanced visual-spatial reasoning capabilities, such as transportation hubs, large-scale retail spaces, urban streets, and public institutions. Our experiments first reveal that even top-tier proprietary models falter, achieving only ~30% success in object and path search. We then use post-training techniques to enhance the open-source Qwen2.5-VL, increasing its success rate by over threefold for both object search (14.83% to 47.38%) and path search (6.44% to 24.94%). Notably, the lower ceiling of path search reveals its inherent difficulty, which we attribute to the demand for sophisticated spatial commonsense. Our results not only show a promising path forward but also quantify the immense challenge that remains in building MLLM agents that can be seamlessly integrated into everyday human life.

视觉搜索类人智能体空间推理全景图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。