arXiv:2603.15679cs.CRcs.AI2026-03中稿 · ICASSP

针对个性化图像生成的滥用风险,提出上下文感知的安全防护机制。

IdentityGuard: Context-Aware Restriction and Provenance for Personalized Synthesis

  • 仅在涉及个人身份时限制有害内容生成,避免全局过滤
  • 通过专属水印实现概念级溯源,追踪生成内容来源
  • 既防滥用又保模型可用性,适合高安全需求场景

个性化文本到图像模型带来独特安全挑战,现有全局无差别过滤方法难以应对。此类方法为防止滥用,不得不彻底删除特定概念,导致模型整体功能受损。本文提出IDENTITYGUARD,基于安全应与威胁同源的思路,采用条件化限制机制——仅当个性化身份与有害内容结合时才触发拦截;并引入概念特异性水印,实现精准溯源。实验表明,该方法可在保障模型可用性的前提下有效防范滥用,并支持可靠追踪。本工作突破了粗粒度全局过滤的局限,为人工智能安全提供更精准、负责任的新路径。

原文摘要 · Abstract (English)

The nature of personalized text-to-image models poses a unique safety challenge that generic context-blind methods are ill-equipped to handle. Such global filters create a dilemma: to prevent misuse, they are forced to damage the model's broader utility by erasing concepts entirely, causing unacceptable collateral damage.Our work presents a more precisely targeted approach, built on the principle that security should be as context-aware as the threat itself, intrinsically bound to the personalized concept. We present IDENTITYGUARD, which realizes this principle through a conditional restriction that blocks harmful content only when combined with the personalized identity, and a concept-specific watermark for precise traceability. Experiments show our approach prevents misuse while preserving the model's utility and enabling robust traceability. By moving beyond blunt, global filters, our work demonstrates a more effective and responsible path toward AI safety.

AI安全图像生成身份保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。