arXiv:2603.02816cs.CVcs.AI2026-03被引 2

让AI视频自动植入品牌,保持原意自然真实

BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation

论文配图:BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
图 1 · 摘自论文原文
  • 分离式多智能体框架,离线建品牌知识库,线上实时优化提示
  • 18个品牌测试中,品牌识别率和语义保真度均显著优于基线
  • 适合广告商、内容创作者,推动AI视频商业化落地

文本生成视频(T2V)模型的快速发展彻底改变了内容创作方式,但其商业潜力尚未被充分挖掘。本文首次提出无缝品牌整合任务:在用户提示生成的视频中自动嵌入广告品牌,同时保持与用户意图一致的语义完整性。该任务面临三大挑战:保持提示忠实性、确保品牌可识别性、实现上下文自然融合。为此,我们提出BrandFusion——一种由两个协同阶段组成的多智能体框架。离线阶段(面向广告商)通过探测模型先验并采用轻量微调适应新品牌,构建品牌知识库;在线阶段(面向用户)由五个智能体协作,通过迭代优化用户提示,利用共享知识库与实时上下文追踪,确保品牌可见性与语义一致性。在18个现成品牌及2个自定义品牌上,跨多个主流T2V模型的实验表明,BrandFusion在语义保留、品牌可识别性和融合自然性方面显著优于基线。人工评估进一步证实用户满意度更高,为可持续的T2V商业化提供了可行路径。

原文摘要 · Abstract (English)

The rapid advancement of text-to-video (T2V) models has revolutionized content creation, yet their commercial potential remains largely untapped. We introduce, for the first time, the task of seamless brand integration in T2V: automatically embedding advertiser brands into prompt-generated videos while preserving semantic fidelity to user intent. This task confronts three core challenges: maintaining prompt fidelity, ensuring brand recognizability, and achieving contextually natural integration. To address them, we propose BrandFusion, a novel multi-agent framework comprising two synergistic phases. In the offline phase (advertiser-facing), we construct a Brand Knowledge Base by probing model priors and adapting to novel brands via lightweight fine-tuning. In the online phase (user-facing), five agents jointly refine user prompts through iterative refinement, leveraging the shared knowledge base and real-time contextual tracking to ensure brand visibility and semantic alignment. Experiments on 18 established and 2 custom brands across multiple state-of-the-art T2V models demonstrate that BrandFusion significantly outperforms baselines in semantic preservation, brand recognizability, and integration naturalness. Human evaluations further confirm higher user satisfaction, establishing a practical pathway for sustainable T2V monetization.

文本生成视频品牌植入多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。