arXiv:2509.15803cs.CVcs.AI2025-09

解决文本生成图像模型过度依赖品牌的问题

CIDER: A Causal Cure for Brand-Obsessed Text-to-Image Models

  • 通过轻量级检测器识别品牌内容,用视觉语言模型生成风格替代方案
  • 在主流模型上测试,显著降低显性和隐性品牌偏见,保持图像质量
  • 无需重训练,适合需快速部署的AI内容生成场景

文本到图像(T2I)模型存在显著但未被充分研究的“品牌偏见”,即在通用提示下倾向于生成主流商业品牌的图像,带来伦理和法律风险。本文提出CIDER,一种无需模型修改的推理时缓解框架,通过提示优化实现低成本干预。CIDER使用轻量级检测器识别品牌内容,并借助视觉语言模型(VLM)生成风格差异化的替代图像。我们引入品牌中立性评分(BNS)量化该问题,并在多个领先T2I模型上进行广泛实验。结果表明,CIDER能有效降低显性和隐性偏见,同时维持图像质量和美学吸引力。本工作为生成更原创、更公平的内容提供了实用方案,助力可信生成式AI的发展。

原文摘要 · Abstract (English)

Text-to-image (T2I) models exhibit a significant yet under-explored "brand bias", a tendency to generate contents featuring dominant commercial brands from generic prompts, posing ethical and legal risks. We propose CIDER, a novel, model-agnostic framework to mitigate bias at inference-time through prompt refinement to avoid costly retraining. CIDER uses a lightweight detector to identify branded content and a Vision-Language Model (VLM) to generate stylistically divergent alternatives. We introduce the Brand Neutrality Score (BNS) to quantify this issue and perform extensive experiments on leading T2I models. Results show CIDER significantly reduces both explicit and implicit biases while maintaining image quality and aesthetic appeal. Our work offers a practical solution for more original and equitable content, contributing to the development of trustworthy generative AI.

图像生成品牌偏见提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。