arXiv:2603.13070cs.CV2026-03被引 1

防止文生图模型记忆训练图像,提升生成多样性与版权安全性

Mitigating Memorization in Text-to-Image Diffusion via Region-Aware Prompt Augmentation and Multimodal Copy Detection

  • 用物体检测定位关键区域,生成语义一致的提示变体增强多样性
  • 融合局部块、全局语义和纹理特征,无需大量标注即可检测复制行为
  • 在保持图像质量前提下降低过拟合,适合注重版权合规的生成应用

当前先进的文生图扩散模型虽能生成高质量图像,但可能记忆并复现训练数据中的图像,带来版权与隐私风险。现有推理时的提示扰动方法(如随机插入标记或嵌入噪声)虽能降低复制率,但常损害图像与提示的一致性及整体质量。为此,本文提出两种互补方法:首先,区域感知提示增强(RAPTA)利用物体检测识别显著区域,生成语义对齐的提示变体,并在训练中随机采样以增加多样性;其次,注意力驱动多模态复制检测(ADMCD)通过轻量级Transformer聚合局部块、全局语义和纹理线索,生成融合表征,并采用简单阈值决策规则检测复制,无需依赖大规模标注数据。实验表明,RAPTA有效缓解过拟合且保持高合成质量,而ADMCD在复制检测上优于单一模态指标。

原文摘要 · Abstract (English)

State-of-the-art text-to-image diffusion models can produce impressive visuals but may memorize and reproduce training images, creating copyright and privacy risks. Existing prompt perturbations applied at inference time, such as random token insertion or embedding noise, may lower copying but often harm image-prompt alignment and overall fidelity. To address this, we introduce two complementary methods. First, Region-Aware Prompt Augmentation (RAPTA) uses an object detector to find salient regions and turn them into semantically grounded prompt variants, which are randomly sampled during training to increase diversity, while maintaining semantic alignment. Second, Attention-Driven Multimodal Copy Detection (ADMCD) aggregates local patch, global semantic, and texture cues with a lightweight transformer to produce a fused representation, and applies simple thresholded decision rules to detect copying without training with large annotated datasets. Experiments show that RAPTA reduces overfitting while maintaining high synthesis quality, and that ADMCD reliably detects copying, outperforming single-modal metrics.

文生图版权安全扩散模型多模态检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。