arXiv:2410.13585cs.CV2024-10中稿 · VCIP 2024被引 2

用普通视频生成伪标签数据,提升跨场景摄像机推荐准确率

Pseudo Dataset Generation for Out-of-Domain Multi-Camera View Recommendation

  • 从普通视频中重构伪多视角标签数据
  • 目标域准确率提升68%相对效果
  • 适合跨场景影视剪辑推荐任务

多摄像机系统在电影、电视剧等媒体制作中至关重要。在每个时间点选择合适的摄像机对制作质量与观众偏好有决定性影响。基于学习的视图推荐框架可辅助专业决策,但通常在训练域外表现不佳。标注的多摄像机视图推荐数据集稀缺加剧了这一问题。我们基于一个关键观察:许多视频由原始多摄像机视频剪辑而成,提出将普通视频转换为伪标注的多摄像机视图推荐数据集。令人振奋的是,仅使用目标域视频生成的伪标签数据训练模型,即可在目标域实现68%的相对准确率提升,并弥合了域内与从未见过域之间的准确率差距。

原文摘要 · Abstract (English)

Multi-camera systems are indispensable in movies, TV shows, and other media. Selecting the appropriate camera at every timestamp has a decisive impact on production quality and audience preferences. Learning-based view recommendation frameworks can assist professionals in decision-making. However, they often struggle outside of their training domains. The scarcity of labeled multi-camera view recommendation datasets exacerbates the issue. Based on the insight that many videos are edited from the original multi-camera videos, we propose transforming regular videos into pseudo-labeled multi-camera view recommendation datasets. Promisingly, by training the model on pseudo-labeled datasets stemming from videos in the target domain, we achieve a 68% relative improvement in the model's accuracy in the target domain and bridge the accuracy gap between in-domain and never-before-seen domains.

视频推荐伪标签多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。