arXiv:2508.20582cs.IR2025-08被引 3

用多模态大模型自动生成高价值广告摘要,提升抖音搜索广告效果。

SUMMA: A Multimodal Large Language Model for Advertisement Summarization

  • 分两阶段训练:先监督微调,再强化学习优化摘要质量。
  • 线上实验显示广告收入提升1.5%,显著优于基线方法。
  • 适合关注广告检索与推荐系统的工程师和算法研究者。

理解多模态视频广告对提升短视频平台的查询-广告匹配度和相关性排序至关重要,有助于提高广告效果与用户体验。然而,由于高度压缩的视频嵌入依赖,多模态信息的高效利用长期受限。为此,我们提出SUMMA(Summarizing MultiModal Ads),一种自动将视频广告转化为高商业价值内容摘要的多模态模型,从而改善其在抖音搜索广告系统中的可理解性与排序表现。SUMMA通过两阶段训练策略——多模态监督微调后接基于混合奖励机制的强化学习——在包含视频帧及ASR/OCR转录文本的领域数据上进行训练,生成具有商业价值且可解释的摘要。我们将SUMMA生成的摘要集成至生产管道,直接提升候选召回与相关性排序阶段的效果。离线与在线实验均表明其性能显著优于基线,线上结果验证了广告收入统计上显著提升1.5%。本工作建立了一种将多模态信息浓缩为代表性文本的新范式,有效对齐视觉广告内容与用户查询意图,在检索与推荐场景中实现精准匹配。

原文摘要 · Abstract (English)

Understanding multimodal video ads is crucial for improving query-ad matching and relevance ranking on short video platforms, enhancing advertising effectiveness and user experience. However, the effective utilization of multimodal information with high commercial value still largely constrained by reliance on highly compressed video embeddings-has long been inadequate. To address this, we propose SUMMA (the abbreviation of Summarizing MultiModal Ads), a multimodal model that automatically processes video ads into summaries highlighting the content of highest commercial value, thus improving their comprehension and ranking in Douyin search-advertising systems. SUMMA is developed via a two-stage training strategy-multimodal supervised fine-tuning followed by reinforcement learning with a mixed reward mechanism-on domain-specific data containing video frames and ASR/OCR transcripts, generating commercially valuable and explainable summaries. We integrate SUMMA-generated summaries into our production pipeline, directly enhancing the candidate retrieval and relevance ranking stages in real search-advertising systems. Both offline and online experiments show substantial improvements over baselines, with online results indicating a statistically significant 1.5% increase in advertising revenue. Our work establishes a novel paradigm for condensing multimodal information into representative texts, effectively aligning visual ad content with user query intent in retrieval and recommendation scenarios.

广告摘要多模态大模型推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。