用文本引导学习图像不变特征,提升水印在各种变形下的鲁棒性。
Text-Guided Image Invariant Feature Learning for Robust Image Watermarking
- 利用CLIP的文本嵌入作为语义锚点,强制特征在变形下保持一致。
- 在多种变换下提取准确率更高,特征一致性测试中余弦相似度更优。
- 适合需要高鲁棒性水印的场景,如版权保护与内容认证。
确保图像水印在各种变换下的鲁棒性对维护内容完整性至关重要。现有自监督学习(SSL)方法如DINO虽被用于水印,但主要关注通用特征表示,未显式学习不变特征。本文提出一种新型文本引导的不变特征学习框架,用于鲁棒图像水印。该方法利用CLIP的多模态能力,以文本嵌入作为稳定语义锚点,强制特征在各类失真下保持不变。我们在多个数据集上评估了该方法,结果表明其在多种图像变换下均表现出更强鲁棒性。相比现有最优的SSL方法,本模型在特征一致性测试中实现了更高的余弦相似度,并在严重失真条件下显著提升了水印提取准确率。这些结果验证了该方法在为深度学习水印量身定制不变表示方面的有效性。
原文摘要 · Abstract (English)
Ensuring robustness in image watermarking is crucial for and maintaining content integrity under diverse transformations. Recent self-supervised learning (SSL) approaches, such as DINO, have been leveraged for watermarking but primarily focus on general feature representation rather than explicitly learning invariant features. In this work, we propose a novel text-guided invariant feature learning framework for robust image watermarking. Our approach leverages CLIP's multimodal capabilities, using text embeddings as stable semantic anchors to enforce feature invariance under distortions. We evaluate the proposed method across multiple datasets, demonstrating superior robustness against various image transformations. Compared to state-of-the-art SSL methods, our model achieves higher cosine similarity in feature consistency tests and outperforms existing watermarking schemes in extraction accuracy under severe distortions. These results highlight the efficacy of our method in learning invariant representations tailored for robust deep learning-based watermarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。