让插入的广告物更贴合场景,自动匹配品牌Logo。
Toward Intelligent Scene Augmentation for Context-Aware Object Placement and Sponsor-Logo Integration
- 根据场景智能选择并生成合适物体
- 自动检测产品并正确添加品牌标识
- 适合广告设计与数字媒体自动化
智能图像编辑日益依赖计算机视觉、多模态推理和生成模型的进步。尽管视觉语言模型(VLMs)和扩散模型能实现可控图像操作,但现有方法很少保证插入物体在语境上合理。本文提出两个新任务:(1) 上下文感知物体插入,需预测合适物体类别、生成并合理放置于场景中;(2) 赞助商品标志增强,包括检测产品并插入正确品牌标识,即使物品无标或标错。为支持这些任务,我们构建了两个新数据集,包含类别标注、放置区域和赞助-产品标签。
原文摘要 · Abstract (English)
Intelligent image editing increasingly relies on advances in computer vision, multimodal reasoning, and generative modeling. While vision-language models (VLMs) and diffusion models enable guided visual manipulation, existing work rarely ensures that inserted objects are \emph{contextually appropriate}. We introduce two new tasks for advertising and digital media: (1) \emph{context-aware object insertion}, which requires predicting suitable object categories, generating them, and placing them plausibly within the scene; and (2) \emph{sponsor-product logo augmentation}, which involves detecting products and inserting correct brand logos, even when items are unbranded or incorrectly branded. To support these tasks, we build two new datasets with category annotations, placement regions, and sponsor-product labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。