arXiv:2411.08753cs.CV2024-11CVPR被引 7

用视频旁白自动选出最清晰的视角,无需人工标注。

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

  • 用视图相关字幕预测准确率作为信息量指标,生成伪标签。
  • 在两个数据集上超越现有方法,人类评估也更优。
  • 适合做多视角教学视频的自动视角推荐系统。

给定一个多视角视频,哪个视角对人类观察者最有益?现有方法依赖启发式规则或昂贵的“最佳视角”标注,限制了应用范围。本文提出一种弱监督方法,利用教学类多视角视频的伴随语言来恢复最具信息量的视角。核心假设是:一个视角能越准确地预测与视角无关的文本摘要,其信息量就越高。为此,我们提出LangView框架,通过视图相关字幕预测的相对准确性,构建最佳视角伪标签,并以此训练视角选择器,同时引入辅助相机位姿预测器增强视角敏感性。推理时仅需输入多视角视频,无需语言或相机位姿,即可输出每个时刻的最佳观看视角。在包含多样化多摄像机布局和操作类活动的两个挑战性数据集上,模型在定量指标和人类评估中均持续优于现有最佳基线。

原文摘要 · Abstract (English)

Given a multi-view video, which viewpoint is most informative for a human observer? Existing methods rely on heuristics or expensive "best-view" supervision to answer this question, limiting their applicability. We propose a weakly supervised approach that leverages language accompanying an instructional multi-view video as a means to recover its most informative viewpoint(s). Our key hypothesis is that the more accurately an individual view can predict a view-agnostic text summary, the more informative it is. To put this into action, we propose LangView, a framework that uses the relative accuracy of view-dependent caption predictions as a proxy for best view pseudo-labels. Then, those pseudo-labels are used to train a view selector, together with an auxiliary camera pose predictor that enhances view-sensitivity. During inference, our model takes as input only a multi-view video--no language or camera poses--and returns the best viewpoint to watch at each timestep. On two challenging datasets comprised of diverse multi-camera setups and how-to activities, our model consistently outperforms state-of-the-art baselines, both with quantitative metrics and human evaluation. Project page: https://vision.cs.utexas.edu/projects/which-view-shows-it-best.

多视角弱监督视频理解语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。