arXiv:2510.12749cs.CV2025-10中稿 · IEEE Transactions …

统一视觉里程计、分割与渲染,实现城市场景的全面理解。

SPORTS: Simultaneous Panoptic Odometry, Rendering, Tracking and Segmentation for Urban Scenes Understanding

  • 用自适应注意力融合多模态特征,提升跨帧对齐精度。
  • 在三个数据集上同时超越现有方法,在定位、分割和新视角生成上表现优异。
  • 适合需要端到端场景理解的自动驾驶与机器人研究者。

场景感知、理解与模拟是具身人工智能智能体的核心技术,但现有方案仍存在分割不全、动态物体干扰、传感器数据稀疏及视图受限等问题。本文提出SPORTS框架,通过将视频全景分割(VPS)、视觉里程计(VO)与场景渲染(SR)任务紧密集成,实现迭代统一的视角理解。VPS采用基于注意力的几何融合机制,利用位姿、深度与光流模态对齐跨帧特征,并设计后匹配策略提升目标身份跟踪效果。在VO中,结合VPS的全景分割结果与光流图,提升动态物体置信度估计,从而增强相机位姿估计精度与深度图生成完整性。此外,点云渲染得益于VO结果,将稀疏点云转换为神经场,合成高保真RGB视图与双全景视图。在三个公开数据集上的大量实验表明,所提注意力特征融合方法在里程计、跟踪、分割与新视角合成任务上均优于多数现有最先进方法。

原文摘要 · Abstract (English)

The scene perception, understanding, and simulation are fundamental techniques for embodied-AI agents, while existing solutions are still prone to segmentation deficiency, dynamic objects' interference, sensor data sparsity, and view-limitation problems. This paper proposes a novel framework, named SPORTS, for holistic scene understanding via tightly integrating Video Panoptic Segmentation (VPS), Visual Odometry (VO), and Scene Rendering (SR) tasks into an iterative and unified perspective. Firstly, VPS designs an adaptive attention-based geometric fusion mechanism to align cross-frame features via enrolling the pose, depth, and optical flow modality, which automatically adjust feature maps for different decoding stages. And a post-matching strategy is integrated to improve identities tracking. In VO, panoptic segmentation results from VPS are combined with the optical flow map to improve the confidence estimation of dynamic objects, which enhances the accuracy of the camera pose estimation and completeness of the depth map generation via the learning-based paradigm. Furthermore, the point-based rendering of SR is beneficial from VO, transforming sparse point clouds into neural fields to synthesize high-fidelity RGB views and twin panoptic views. Extensive experiments on three public datasets demonstrate that our attention-based feature fusion outperforms most existing state-of-the-art methods on the odometry, tracking, segmentation, and novel view synthesis tasks.

全景分割视觉里程计场景渲染端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。