arXiv:2602.22299cs.MMcs.AI2026-02被引 1

用多模态大模型分析视频广告前3秒,提升广告吸引力预测能力。

Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video Ads

  • 基于Transformer的多模态大模型,融合视觉、音频与文本特征
  • 前3秒关键帧分析显示与转化率存在显著相关性
  • 适合广告优化与数字营销团队快速评估广告开头效果

视频广告是品牌吸引消费者的重要媒介,社交平台利用用户数据优化投放以提升参与度。一个关键但研究不足的环节是‘钩子期’——前3秒内捕捉注意力并影响互动表现的关键窗口。由于视频内容包含视觉、听觉和文本等多模态元素,传统方法难以捕捉其复杂交互,亟需先进分析框架。本文提出一种基于变压器架构的多模态大语言模型(MLLM)框架,用于分析视频广告的钩子期。采用均匀随机采样与关键帧选择两种策略,确保声学特征提取的平衡性与代表性。通过先进的MLLM对钩子片段进行描述性分析,并利用BERTopic将输出提炼为高层抽象主题。框架还整合音频属性与聚合广告定向信息,丰富特征体系。在社交媒体平台大规模真实数据上的实证验证表明,该框架有效揭示了钩子期特征与转化率/投资回报率等核心指标间的关联。结果证明该方法具备实用价值与预测能力,为优化视频广告策略提供有力支持。本研究推动了视频广告分析的发展,提供了可扩展的初始时刻理解与优化方法。

原文摘要 · Abstract (English)

Video-based ads are a vital medium for brands to engage consumers, with social media platforms leveraging user data to optimize ad delivery and boost engagement. A crucial but under-explored aspect is the 'hooking period', the first three seconds that capture viewer attention and influence engagement metrics. Analyzing this brief window is challenging due to the multimodal nature of video content, which blends visual, auditory, and textual elements. Traditional methods often miss the nuanced interplay of these components, requiring advanced frameworks for thorough evaluation. This study presents a framework using transformer-based multimodal large language models (MLLMs) to analyze the hooking period of video ads. It tests two frame sampling strategies, uniform random sampling and key frame selection, to ensure balanced and representative acoustic feature extraction, capturing the full range of design elements. The hooking video is processed by state-of-the-art MLLMs to generate descriptive analyses of the ad's initial impact, which are distilled into coherent topics using BERTopic for high-level abstraction. The framework also integrates features such as audio attributes and aggregated ad targeting information, enriching the feature set for further analysis. Empirical validation on large-scale real-world data from social media platforms demonstrates the efficacy of our framework, revealing correlations between hooking period features and key performance metrics like conversion per investment. The results highlight the practical applicability and predictive power of the approach, offering valuable insights for optimizing video ad strategies. This study advances video ad analysis by providing a scalable methodology for understanding and enhancing the initial moments of video advertisements.

视频广告多模态大模型广告优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。