arXiv:2607.00189cs.CV2026-07中稿 · ECCV

利用视频编码器信息提升压缩流下的视觉里程计精度

VOCA: Visual Odometry with Codec Awareness

论文配图:VOCA: Visual Odometry with Codec Awareness
图 1 · 摘自论文原文
  • 引入编码器元信息,增强对压缩图像失真的鲁棒性
  • 在压缩视频流上实现最低相对轨迹误差与绝对误差
  • 适合部署在真实硬件压缩场景的自动驾驶系统

从图像流中估计相机位姿是将感知融入规划与决策的空间世界模型的关键。几乎所有视觉里程计(VO)和同步定位与地图构建(V-SLAM)系统均基于原始无损视频数据集训练。然而,实际应用中普遍使用硬件编码解码器对视频流进行高效压缩,显著降低存储与带宽开销。这种有损压缩引入视觉伪影,严重影响传统追踪系统的性能。本文提出VOCA,一种利用编码器信息的因果立体视觉里程计方法,显著提升压缩流上的追踪表现。在压缩视频流上,该方法在相对轨迹误差、效率和绝对轨迹误差三项指标上达到当前最优水平。本工作揭示了利用广泛存在的视频编码器信息提升视觉任务潜力。

原文摘要 · Abstract (English)

Camera pose estimation from image streams is a critical component of spatial world models that integrate perception into planning and decision-making. Nearly all Visual Odometry (VO) and Simultaneous Localization and Mapping (V-SLAM) systems have focused on datasets containing raw, uncompressed videos. Many working systems instead use ubiquitous hardware units to efficiently compress and decode video streams, saving orders of magnitude in storage and bandwidth. However, this lossy compression introduces visual artifacts that hinder the performance of traditional tracking systems. We present VOCA, a causal stereo visual-odometry method that exploits codec information to improve tracking performance. We achieve state-of-the-art performance on causal VO for relative trajectory error, efficiency, and absolute trajectory error on compressed streams. This work highlights the potential of leveraging widely available video codec information for vision tasks.

视觉里程计视频压缩编码器信息三维重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。