arXiv:2507.07056cs.CRcs.LG2025-07KDD被引 2

为个性化LoRA模型提供免数据的安全编辑,防止恶意内容生成。

LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing

  • 通过对抗优化与语义增强动态重置LoRA权重空间。
  • 在不损失正常功能前提下,有效阻断恶意生成内容。
  • 适合平台方部署,保障个性化模型共享安全。

低秩适配(LoRA)模型的普及使用户能轻松分享轻量级个性化文本到图像生成模型(如个人肖像),在Civitai、Liblib等平台广泛传播。然而,这种‘共享即使用’生态带来严重风险:良性LoRA可能被攻击者用于生成政治、诽谤性图像等有害内容,侵害创作者权益并威胁平台安全。现有防御方法聚焦于完整扩散模型,忽视了LoRA作为模块化适配器的独特角色及其对对抗提示工程的脆弱性。为此,我们提出LoRAShield,首个无需数据的免数据编辑框架,用于保护LoRA模型免遭滥用。该平台驱动方法通过对抗优化和语义增强,动态编辑并重新对齐LoRA的权重子空间。实验表明,LoRAShield在阻止恶意生成方面展现出卓越的有效性、效率与鲁棒性,同时不损害原有良性任务的功能。将防御机制迁移至平台端,推动了个性化模型安全、可扩展的共享,是构建可信生成生态的关键一步。

原文摘要 · Abstract (English)

The proliferation of Low-Rank Adaptation (LoRA) models has democratized personalized text-to-image generation, enabling users to share lightweight models (e.g., personal portraits) on platforms like Civitai and Liblib. However, this "share-and-play" ecosystem introduces critical risks: benign LoRAs can be weaponized by adversaries to generate harmful content (e.g., political, defamatory imagery), undermining creator rights and platform safety. Existing defenses like concept-erasure methods focus on full diffusion models (DMs), neglecting LoRA's unique role as a modular adapter and its vulnerability to adversarial prompt engineering. To bridge this gap, we propose LoRAShield, the first data-free editing framework for securing LoRA models against misuse. Our platform-driven approach dynamically edits and realigns LoRA's weight subspace via adversarial optimization and semantic augmentation. Experimental results demonstrate that LoRAShield achieves remarkable effectiveness, efficiency, and robustness in blocking malicious generations without sacrificing the functionality of the benign task. By shifting the defense to platforms, LoRAShield enables secure, scalable sharing of personalized models, a critical step toward trustworthy generative ecosystems.

LoRA安全生成模型内容防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。