为自回归图像生成设计安全水印框架,兼顾画质与防伪能力。
Safe-VAR: Safe Visual Autoregressive Model for Text-to-Image Generative Watermarking
- 根据水印复杂度动态选择嵌入时机,提升鲁棒性。
- 多尺度融合机制有效处理内容与水印的交互关系。
- 在画质、水印保真度和抗干扰性上均领先现有方法。
随着自回归学习在大语言模型中的成功,其已成为文本到图像生成的主流方法,具备高效性和高质量输出。然而,针对视觉自回归(VAR)模型的不可见水印技术仍缺乏研究,而这一技术对防止滥用至关重要。现有水印方法主要面向扩散模型,难以适应VAR模型的序列生成特性。为此,我们提出Safe-VAR,首个专为自回归文本到图像生成设计的水印框架。研究发现,水印注入时机显著影响生成质量,不同复杂度的水印存在最佳注入时间。基于此,我们提出自适应尺度交互模块,根据水印信息与图像视觉特征动态选择最优嵌入策略,确保水印鲁棒性的同时最小化对画质的影响。此外,引入跨尺度融合机制,结合多头与专家混合结构,有效融合多分辨率特征,处理图像内容与水印模式间的复杂交互。实验表明,Safe-VAR在图像质量、水印保真度及抗扰动鲁棒性方面均达到当前最优水平。此外,该方法对域外水印数据集二维码表现出强泛化能力。
原文摘要 · Abstract (English)
With the success of autoregressive learning in large language models, it has become a dominant approach for text-to-image generation, offering high efficiency and visual quality. However, invisible watermarking for visual autoregressive (VAR) models remains underexplored, despite its importance in misuse prevention. Existing watermarking methods, designed for diffusion models, often struggle to adapt to the sequential nature of VAR models. To bridge this gap, we propose Safe-VAR, the first watermarking framework specifically designed for autoregressive text-to-image generation. Our study reveals that the timing of watermark injection significantly impacts generation quality, and watermarks of different complexities exhibit varying optimal injection times. Motivated by this observation, we propose an Adaptive Scale Interaction Module, which dynamically determines the optimal watermark embedding strategy based on the watermark information and the visual characteristics of the generated image. This ensures watermark robustness while minimizing its impact on image quality. Furthermore, we introduce a Cross-Scale Fusion mechanism, which integrates mixture of both heads and experts to effectively fuse multi-resolution features and handle complex interactions between image content and watermark patterns. Experimental results demonstrate that Safe-VAR achieves state-of-the-art performance, significantly surpassing existing counterparts regarding image quality, watermarking fidelity, and robustness against perturbations. Moreover, our method exhibits strong generalization to an out-of-domain watermark dataset QR Codes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。