用知识库提升广告视频生成的语义对齐与动作自然度
KD-CVG: A Knowledge-Driven Approach for Creative Video Generation

- 构建广告创意知识库,通过语义感知检索增强理解
- 在多个基准上实现更优的语义对齐和动作适应性
- 适合广告生成、视频内容创作等场景的研究者使用
创意生成(CG)利用生成模型自动产出突出产品特点的广告内容,是当前研究热点。然而,尽管生成文本和图像的进展显著,广告视频生成(CVG)仍相对未被充分探索。这主要源于文本到视频(T2V)模型面临的两大挑战:(a)语义对齐模糊,难以准确关联产品卖点与视频内容;(b)运动适应性不足,导致动作不自然或失真。为此,我们构建了广告创意知识库(ACKB),提出一种知识驱动方法(KD-CVG)以克服现有模型的知识局限。KD-CVG包含两个核心模块:语义感知检索(SAR)和多模态知识参考(MKR)。SAR利用图注意力网络的语义感知能力与强化学习反馈,提升模型对卖点与视频关联的理解;在此基础上,MKR将语义和运动先验融入T2V模型,弥补知识空白。大量实验验证了该方法在语义对齐与运动适应性方面的优越性能,优于现有先进方法。代码与数据集将开源于 https://kdcvg.github.io/KDCVG/。
原文摘要 · Abstract (English)
Creative Generation (CG) leverages generative models to automatically produce advertising content that highlights product features, and it has been a significant focus of recent research. However, while CG has advanced considerably, most efforts have concentrated on generating advertising text and images, leaving Creative Video Generation (CVG) relatively underexplored. This gap is largely due to two major challenges faced by Text-to-Video (T2V) models: (a) \textbf{ambiguous semantic alignment}, where models struggle to accurately correlate product selling points with creative video content, and (b) \textbf{inadequate motion adaptability}, resulting in unrealistic movements and distortions. To address these challenges, we develop a comprehensive Advertising Creative Knowledge Base (ACKB) as a foundational resource and propose a knowledge-driven approach (KD-CVG) to overcome the knowledge limitations of existing models. KD-CVG consists of two primary modules: Semantic-Aware Retrieval (SAR) and Multimodal Knowledge Reference (MKR). SAR utilizes the semantic awareness of graph attention networks and reinforcement learning feedback to enhance the model's comprehension of the connections between selling points and creative videos. Building on this, MKR incorporates semantic and motion priors into the T2V model to address existing knowledge gaps. Extensive experiments have demonstrated KD-CVG's superior performance in achieving semantic alignment and motion adaptability, validating its effectiveness over other state-of-the-art methods. The code and dataset will be open source at https://kdcvg.github.io/KDCVG/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。