arXiv:2503.17975cs.CVcs.AI2025-03被引 5

解决短视频剪辑中镜头顺序排列难题,提升视频叙事效果。

Shot Sequence Ordering for Video Editing: Benchmarks, Metrics, and Cinematology-Inspired Computing Methods

  • 引入肯德尔τ距离作为排序评估指标,设计新损失函数。
  • 构建两个公开数据集,准确率显著优于基线模型。
  • 融合电影元数据与镜头标签,适合影视创作与AI编辑研究者。

随着短视频平台兴起,视频制作需求激增,但高质量内容仍依赖专业剪辑技巧和视觉语言理解。为推动这一领域发展,本文提出镜头序列排序(SSO)任务,并针对缺乏公开基准数据的瓶颈,构建了两个新数据集:AVE-Order和ActivityNet-Order。采用肯德尔τ距离作为评估指标,提出肯德尔τ距离-交叉熵损失函数。同时引入电影学嵌入(Cinematology Embedding),将电影元数据和镜头标签作为先验知识融入模型,构建AVE-Meta数据集验证有效性。实验表明,所提方法在排序准确率上显著提升。所有数据集已开源:https://github.com/litchiar/ShotSeqBench。

原文摘要 · Abstract (English)

With the rising popularity of short video platforms, the demand for video production has increased substantially. However, high-quality video creation continues to rely heavily on professional editing skills and a nuanced understanding of visual language. To address this challenge, the Shot Sequence Ordering (SSO) task in AI-assisted video editing has emerged as a pivotal approach for enhancing video storytelling and the overall viewing experience. Nevertheless, the progress in this field has been impeded by a lack of publicly available benchmark datasets. In response, this paper introduces two novel benchmark datasets, AVE-Order and ActivityNet-Order. Additionally, we employ the Kendall Tau distance as an evaluation metric for the SSO task and propose the Kendall Tau Distance-Cross Entropy Loss. We further introduce the concept of Cinematology Embedding, which incorporates movie metadata and shot labels as prior knowledge into the SSO model, and constructs the AVE-Meta dataset to validate the method's effectiveness. Experimental results indicate that the proposed loss function and method substantially enhance SSO task accuracy. All datasets are publicly accessible at https://github.com/litchiar/ShotSeqBench.

视频剪辑镜头排序电影学嵌入数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。