提出通用视频同步框架VideoSync,不依赖音频或特定视觉特征。
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
- 设计独立于特征提取方法的同步框架,适配多种内容类型。
- 在新数据集上验证,纠正旧方法偏差后性能显著提升。
- 适合需要跨视角视频对齐的现实场景,如安防、体育分析。
视频同步——将不同角度拍摄同一事件的多路视频对齐——对实景电视制作、体育分析、监控及自动驾驶系统至关重要。以往工作严重依赖音频信号或特定视觉事件,导致在信号不可靠或缺失的场景下适用性受限。此外,现有基准缺乏通用性与可复现性,阻碍了领域进展。本文提出VideoSync框架,不依赖特定特征提取方法(如人体姿态估计),实现更广泛的应用覆盖。我们在涵盖单人、多人及非人类场景的新数据集上评估系统,并提供数据构建方法与代码,建立可复现的基准。分析发现,先前最先进方法SeSyn-Net存在预处理偏差,导致性能高估。我们修正该偏差,提出更严格的评估框架,证明VideoSync在公平条件下优于现有方法。同时,探索多种同步偏移预测方法,确认基于卷积神经网络(CNN)的模型表现最佳。研究推动视频同步摆脱领域限制,更具泛化性与鲁棒性。
原文摘要 · Abstract (English)
Video synchronization-aligning multiple video streams capturing the same event from different angles-is crucial for applications such as reality TV show production, sports analysis, surveillance, and autonomous systems. Prior work has heavily relied on audio cues or specific visual events, limiting applicability in diverse settings where such signals may be unreliable or absent. Additionally, existing benchmarks for video synchronization lack generality and reproducibility, restricting progress in the field. In this work, we introduce VideoSync, a video synchronization framework that operates independently of specific feature extraction methods, such as human pose estimation, enabling broader applicability across different content types. We evaluate our system on newly composed datasets covering single-human, multi-human, and non-human scenarios, providing both the methodology and code for dataset creation to establish reproducible benchmarks. Our analysis reveals biases in prior SOTA work, particularly in SeSyn-Net's preprocessing pipeline, leading to inflated performance claims. We correct these biases and propose a more rigorous evaluation framework, demonstrating that VideoSync outperforms existing approaches, including SeSyn-Net, under fair experimental conditions. Additionally, we explore various synchronization offset prediction methods, identifying a convolutional neural network (CNN)-based model as the most effective. Our findings advance video synchronization beyond domain-specific constraints, making it more generalizable and robust for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。