用环境音辅助视频相机位姿估计,提升视觉失效时的鲁棒性。
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
- 融合声音方向谱与双耳嵌入特征,增强视觉位姿模型
- 在两个真实数据集上超越纯视觉基线,视觉受损时仍稳定
- 首次实现真实视频中利用日常声音做相机位姿估计
理解相机运动是具身感知与三维场景理解的基础问题。尽管视觉方法发展迅速,但在运动模糊或遮挡等视觉退化条件下常表现不佳。本文表明,被动采集的环境声音可为真实场景视频中的相对相机位姿估计提供互补信息。我们提出一种简单有效的音视频框架,将到达方向(DOA)频谱与双耳化嵌入特征融入先进的纯视觉位姿估计模型。在两个大规模数据集上的实验显示,该方法持续优于强视觉基线,且在视觉信息受损时仍保持鲁棒性。据我们所知,这是首个成功在真实视频中利用音频进行相对相机位姿估计的工作,证实了日常偶然声音作为经典空间问题的潜在信号。项目主页:http://vision.cs.utexas.edu/projects/av_camera_pose。
原文摘要 · Abstract (English)
Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly, they often struggle under visually degraded conditions such as motion blur or occlusions. In this work, we show that passive scene sounds provide cues complementary to vision for relative camera pose estimation for in-the-wild videos. We introduce a simple but effective audio-visual framework that integrates direction-of-arrival (DOA) spectra and binauralized embeddings into a state-of-the-art vision-only pose estimation model. Our results on two large datasets show consistent gains over strong visual baselines, plus robustness when the visual information is corrupted. To our knowledge, this represents the first work to successfully leverage audio for relative camera pose estimation in real-world videos, and it establishes incidental, everyday audio as an unexpected but promising signal for a classic spatial challenge. Project: http://vision.cs.utexas.edu/projects/av_camera_pose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。