用学习到的连贯性评分自动组装多镜头视频,减少对文字依赖。
SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
- 基于学习的片段组装评分衡量镜头间时序与语义连贯性
- 在两个数据集上比现有方法提升48.6%的IoU并加速43%
- 适合无需文字描述的自动化视频生成场景
我们提出SKALD,一种多镜头视频组装方法,从候选镜头中构建连贯视频序列,尽量减少对文本的依赖。核心是学习型片段组装(LCA)评分,通过学习模型衡量镜头间的时序与语义关系,量化叙事连贯性。针对多镜头组合的指数级复杂度,采用基于LCA评分的高效束搜索算法。为在有限人工标注下有效训练,提出两项任务:镜头连贯性学习(使用对比学习区分连贯与不连贯序列)和特征回归(将学习表示转为实值连贯性分数)。开发两种变体:仅依赖视觉连贯性的基础版SKALD,以及可整合辅助文本信息的SKALD-text。在VSPD和自建的MSV3C数据集上的实验表明,SKALD在IoU上相比最优方法最高提升48.6%,速度加快43%。用户研究进一步验证,45%参与者更偏好SKALD生成视频,高于文本驱动方法的22%。
原文摘要 · Abstract (English)
We present SKALD, a multi-shot video assembly method that constructs coherent video sequences from candidate shots with minimal reliance on text. Central to our approach is the Learned Clip Assembly (LCA) score, a learning-based metric that measures temporal and semantic relationships between shots to quantify narrative coherence. We tackle the exponential complexity of combining multiple shots with an efficient beam-search algorithm guided by the LCA score. To train our model effectively with limited human annotations, we propose two tasks for the LCA encoder: Shot Coherence Learning, which uses contrastive learning to distinguish coherent and incoherent sequences, and Feature Regression, which converts these learned representations into a real-valued coherence score. We develop two variants: a base SKALD model that relies solely on visual coherence and SKALD-text, which integrates auxiliary text information when available. Experiments on the VSPD and our curated MSV3C datasets show that SKALD achieves an improvement of up to 48.6% in IoU and a 43% speedup over the state-of-the-art methods. A user study further validates our approach, with 45% of participants favoring SKALD-assembled videos, compared to 22% preferring text-based assembly methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。