用物体运动匹配实现多摄像头视频毫秒级同步。
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
- 基于跨视角物体运动的几何约束,自动对齐多摄像头视频。
- 在4个数据集上中位误差低于50毫秒,优于现有方法。
- 无需特殊设备或人工标注,适合家庭聚会等场景。
如今,人们可用多个消费级摄像机记录音乐会、体育赛事、讲座、家庭聚会和生日派对等难忘时刻。然而,多摄像头视频流的同步仍具挑战性。现有方法依赖受控环境、特定目标、人工校正或昂贵硬件。本文提出VisualSync,一种基于多视角动态的优化框架,可在无标定、未同步的视频中实现毫秒级精确对齐。核心洞察是:当一个三维点在两台相机中同时可见时,一旦时间对齐,其运动将满足对极几何约束。为此,VisualSync利用现成的3D重建、特征匹配与稠密跟踪技术提取轨迹片段、相对位姿和跨视图对应关系,联合最小化对极误差以估计各相机的时间偏移。在四个多样且具有挑战性的数据集上的实验表明,VisualSync显著优于基线方法,中位同步误差低于50毫秒。
原文摘要 · Abstract (English)
Today, people can easily record memorable moments, ranging from concerts, sports events, lectures, family gatherings, and birthday parties with multiple consumer cameras. However, synchronizing these cross-camera streams remains challenging. Existing methods assume controlled settings, specific targets, manual correction, or costly hardware. We present VisualSync, an optimization framework based on multi-view dynamics that aligns unposed, unsynchronized videos at millisecond accuracy. Our key insight is that any moving 3D point, when co-visible in two cameras, obeys epipolar constraints once properly synchronized. To exploit this, VisualSync leverages off-the-shelf 3D reconstruction, feature matching, and dense tracking to extract tracklets, relative poses, and cross-view correspondences. It then jointly minimizes the epipolar error to estimate each camera's time offset. Experiments on four diverse, challenging datasets show that VisualSync outperforms baseline methods, achieving an median synchronization error below 50 ms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。