ContextIQ通过多模态专家系统实现精准视频广告匹配。
ContextIQ: A Multimodal Expert-Based Video Retrieval System for Contextual Advertising
- 采用视频、音频、字幕、元数据等多模态专家分别建模。
- 无需联合训练,在多个基准上达到或超越顶尖模型性能。
- 适合需要品牌安全过滤的广告生态,提升投放精准度。
情境化广告根据用户观看内容投放相关广告。社交平台和流媒体上视频内容的快速增长以及隐私问题,推动了情境化广告的需求。在恰当情境下投放合适广告,可提升用户参与度与广告变现效果。技术层面,高效的情境化广告依赖于能细粒度理解复杂视频内容的视频检索系统。现有基于联合多模态训练的文本到视频检索模型需大量数据与算力,实用性受限,且缺乏广告生态集成所需功能。我们提出ContextIQ,一个专为情境化广告设计的多模态专家型视频检索系统。该系统利用视频、音频、字幕(字幕)及对象、动作、情感等元数据等模态专用专家,构建语义丰富的视频表征。实验表明,本系统无需联合训练,在多个文本到视频检索基准上表现优于或相当主流模型与商业方案。消融研究证实,多模态融合显著提升检索精度,优于仅用视觉-语言模型。此外,我们展示了如ContextIQ的视频检索系统如何融入广告生态,同时保障品牌安全,过滤不当内容。
原文摘要 · Abstract (English)
Contextual advertising serves ads that are aligned to the content that the user is viewing. The rapid growth of video content on social platforms and streaming services, along with privacy concerns, has increased the need for contextual advertising. Placing the right ad in the right context creates a seamless and pleasant ad viewing experience, resulting in higher audience engagement and, ultimately, better ad monetization. From a technology standpoint, effective contextual advertising requires a video retrieval system capable of understanding complex video content at a very granular level. Current text-to-video retrieval models based on joint multimodal training demand large datasets and computational resources, limiting their practicality and lacking the key functionalities required for ad ecosystem integration. We introduce ContextIQ, a multimodal expert-based video retrieval system designed specifically for contextual advertising. ContextIQ utilizes modality-specific experts-video, audio, transcript (captions), and metadata such as objects, actions, emotion, etc.-to create semantically rich video representations. We show that our system, without joint training, achieves better or comparable results to state-of-the-art models and commercial solutions on multiple text-to-video retrieval benchmarks. Our ablation studies highlight the benefits of leveraging multiple modalities for enhanced video retrieval accuracy instead of using a vision-language model alone. Furthermore, we show how video retrieval systems such as ContextIQ can be used for contextual advertising in an ad ecosystem while also addressing concerns related to brand safety and filtering inappropriate content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。