从无标签视频中学习如何自动选择最佳视角,生成更清晰的教程视频。
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
- 通过伪标签训练模型识别视频中的主视角(第一人称/第三人称)。
- 在HowTo100M和Ego-Exo4D数据集上实现精准视点切换,准确率超基准方法。
- 适用于少标注场景,适合视频自动生成与智能剪辑应用。
我们提出SWITCH-A-VIEW,一种从无标签但经人工编辑的视频中学习自动选择教程视频播放视角的模型。核心思路是设计预训练任务,对训练视频片段进行主视角(第一人称或第三人称)的伪标注,并挖掘视频画面内容与语音描述之间与视角切换时刻的关联模式。借助该预测器,模型可推广至新多视角场景,实现无监督或弱监督下的视角动态选择。我们在HowTo100M和Ego-Exo4D等多个真实世界视频数据集上验证了其有效性,显著优于现有基线方法。
原文摘要 · Abstract (English)
We introduce SWITCH-A-VIEW, a model that learns to automatically select the viewpoint to display at each timepoint when creating a how-to video. The key insight of our approach is how to train such a model from unlabeled -- but human-edited -- video samples. We pose a pretext task that pseudo-labels segments in the training videos for their primary viewpoint (egocentric or exocentric), and then discovers the patterns between the visual and spoken content in a how-to video on the one hand and its view-switch moments on the other hand. Armed with this predictor, our model can be applied to new multi-view video settings for orchestrating which viewpoint should be displayed when, even when such settings come with limited labels. We demonstrate our idea on a variety of real-world videos from HowTo100M and Ego-Exo4D, and rigorously validate its advantages. Project: https://vision.cs.utexas.edu/projects/switch_a_view/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。