arXiv:2509.12989cs.CV2025-09被引 12

提出全景视觉系统架构PANORAMA,推动具身智能时代环境感知升级

PANORAMA: The Rise of Omnidirectional Vision in the Embodied AI Era

  • 构建四子系统全景架构PANORAMA,整合生成、感知与理解能力
  • 推动全景视觉在机器人、工业检测等领域的应用落地
  • 适合关注具身智能与全景视觉融合的研究者和工程师

全景视觉通过360度视角理解环境,在机器人、工业检测和环境监测等领域日益重要。相比传统针孔视觉,全景视觉提供更完整的场景感知,显著提升决策可靠性。然而该领域基础研究长期滞后。本次报告揭示具身智能时代下全景视觉的快速发展趋势,由产业需求与学术兴趣共同驱动。重点介绍全景生成、感知、理解等方向的最新突破及配套数据集。基于学界与产业洞察,提出具身智能时代的理想全景系统架构PANORAMA,包含四个关键子系统。同时探讨全景视觉与具身智能交叉领域的新兴趋势、跨社区影响,以及未来路线图与开放挑战。本综述整合前沿进展,阐明构建鲁棒、通用全景智能系统的关键挑战与机遇。

原文摘要 · Abstract (English)

Omnidirectional vision, using 360-degree vision to understand the environment, has become increasingly critical across domains like robotics, industrial inspection, and environmental monitoring. Compared to traditional pinhole vision, omnidirectional vision provides holistic environmental awareness, significantly enhancing the completeness of scene perception and the reliability of decision-making. However, foundational research in this area has historically lagged behind traditional pinhole vision. This talk presents an emerging trend in the embodied AI era: the rapid development of omnidirectional vision, driven by growing industrial demand and academic interest. We highlight recent breakthroughs in omnidirectional generation, omnidirectional perception, omnidirectional understanding, and related datasets. Drawing on insights from both academia and industry, we propose an ideal panoramic system architecture in the embodied AI era, PANORAMA, which consists of four key subsystems. Moreover, we offer in-depth opinions related to emerging trends and cross-community impacts at the intersection of panoramic vision and embodied AI, along with the future roadmap and open challenges. This overview synthesizes state-of-the-art advancements and outlines challenges and opportunities for future research in building robust, general-purpose omnidirectional AI systems in the embodied AI era.

全景视觉具身智能环境感知AI架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。