自动剪辑广告视频,兼顾音画重要性,提升剪辑效率。
AdSum: Two-stream Audio-visual Summarization for Automated Video Advertisement Clipping
- 双流模型融合音画信息,预测帧重要性
- 在真实广告数据集上表现优于现有方法
- 适合广告自动化制作与数字营销从业者
广告商常需为同一广告生成不同长度的多个版本。传统方法依赖人工选取和重剪辑长版广告,耗时费力。本文提出一种基于视频摘要的自动化广告剪辑框架,首次将广告剪辑定义为面向广告场景的镜头选择问题。不同于侧重视觉内容的通用摘要方法,本工作强调音频在广告中的关键作用。为此,构建了双流音画融合模型,预测视频帧被选入官方短版广告的概率。为弥补广告专用数据集缺失,构建了包含102对30秒与15秒广告的真实广告数据集AdSum204。大量实验表明,该模型在平均精度、曲线下面积、斯皮尔曼相关系数和肯德尔等级相关系数等多项指标上均超越现有方法。代码与数据已开源。
原文摘要 · Abstract (English)
Advertisers commonly need multiple versions of the same advertisement (ad) at varying durations for a single campaign. The traditional approach involves manually selecting and re-editing shots from longer video ads to create shorter versions, which is labor-intensive and time-consuming. In this paper, we introduce a framework for automated video ad clipping using video summarization techniques. We are the first to frame video clipping as a shot selection problem, tailored specifically for advertising. Unlike existing general video summarization methods that primarily focus on visual content, our approach emphasizes the critical role of audio in advertising. To achieve this, we develop a two-stream audio-visual fusion model that predicts the importance of video frames, where importance is defined as the likelihood of a frame being selected in the firm-produced short ad. To address the lack of ad-specific datasets, we present AdSum204, a novel dataset comprising 102 pairs of 30-second and 15-second ads from real advertising campaigns. Extensive experiments demonstrate that our model outperforms state-of-the-art methods across various metrics, including Average Precision, Area Under Curve, Spearman, and Kendall. The dataset and code are available at https://github.com/ostadabbas/AdSum204.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。