提出品牌安全生成新任务,消除文字到图像中的商标与设计特征。
From Unlearning to UNBRANDING: A Benchmark for Trademark-Safe Text-to-Image Generation
- 定义'去品牌化'任务,精准移除商标和车辆格栅等细微品牌特征。
- 构建首个包含人类标注的基准数据集,结合视觉语言模型与分割分类器评估。
- 发现高保真模型更易生成品牌标识,凸显该问题的紧迫性与特殊性。
文本到图像扩散模型的快速发展引发了对未经授权复制商标内容的严重担忧。现有研究多聚焦于通用概念(如风格、名人),却未能解决具体品牌标识问题。品牌识别具有多维性,不仅包括明确的标志,还涵盖独特的结构特征(如汽车前格栅)。为此,我们提出去品牌化(unbranding)这一新任务,旨在细粒度地移除商标及微妙的结构性品牌特征,同时保持语义连贯性。我们构建了一个基准数据集,并引入一种新的评估框架,结合视觉语言模型(VLMs)与基于人类标注的标志和商品外观特征的分割分类器,克服了现有品牌检测器无法捕捉抽象商品外观的局限。此外,我们发现更先进的高保真系统(SDXL、FLUX)比旧模型更容易合成品牌标识,凸显该挑战的紧迫性。结果表明,去品牌化是一个需要专门技术的独立问题。
原文摘要 · Abstract (English)
The rapid progress of text-to-image diffusion models raises significant concerns regarding the unauthorized reproduction of trademarked content. While prior work targets general concepts (e.g., styles, celebrities), it fails to address specific brand identifiers. Brand recognition is multi-dimensional, extending beyond explicit logos to encompass distinctive structural features (e.g., a car's front grille). To tackle this, we introduce unbranding, a novel task for the fine-grained removal of both trademarks and subtle structural brand features, while preserving semantic coherence. We construct a benchmark dataset and introduce a novel evaluation framework combining Vision Language Models (VLMs) with segmentation-based classifiers trained on human annotations of logos and trade dress features, addressing the limitations of existing brand detectors that fail to capture abstract trade dress. Furthermore, we observe that newer, higher-fidelity systems (SDXL, FLUX) synthesize brand identifiers more readily than older models, highlighting the urgency of this challenge. Our results confirm that unbranding is a distinct problem requiring specialized techniques. Project Page: https://gmum.github.io/UNBRANDING/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。