用领域知识指导大模型,分析5万条视频广告的故事结构
MLLM-VADStory: Domain Knowledge-Driven Multimodal LLMs for Video Ad Storyline Insights
- 按广告功能拆分片段,用专属分类体系识别每段作用
- 发现故事型创意能提升视频留存率,提炼出高效叙事弧线
- 适合广告主和内容创作者优化视频创意设计
我们提出MLLM-VADStory,一种基于领域知识的多模态大模型框架,可规模化量化与生成视频广告叙事理解的洞察。该框架核心思想是广告叙事由功能意图驱动,每个场景单元承担特定传播功能,在数秒内传递产品与品牌信息。MLLM-VADStory将广告分割为功能单元,利用新型广告专用功能角色分类体系识别其功能,并聚合跨广告的功能序列以还原数据驱动的叙事结构。在四个行业子领域的5万条社交媒体视频广告上应用该框架,发现基于故事的创意能提升视频留存率,并推荐了表现最优的叙事弧线,用于指导广告创作。该框架展示了利用领域知识引导多模态大模型生成可扩展洞察的价值,可广泛应用于视频创意的理解。
原文摘要 · Abstract (English)
We propose MLLM-VADStory, a novel domain knowledge-guided multimodal large language models (MLLM) framework to systematically quantify and generate insights for video ad storyline understanding at scale. The framework is centered on the core idea that ad narratives are structured by functional intent, with each scene unit performing a distinct communicative function, delivering product and brand-oriented information within seconds. MLLM-VADStory segments ads into functional units, classifies each unit's functionality using a novel advertising-specific functional role taxonomy, and then aggregates functional sequences across ads to recover data-driven storyline structures. Applying the framework to 50k social media video ads across four industry subverticals, we find that story-based creatives improve video retention, and we recommend top-performing story arcs to guide advertisers in creative design. Our framework demonstrates the value of using domain knowledge to guide MLLMs in generating scalable insights for video ad storylines, making it a versatile tool for understanding video creatives in general.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。