arXiv:2512.19438cs.CVcs.AI2025-12

通过双向协作提升图像水印嵌入与提取的鲁棒性

MT-Mark: Rethinking Image Watermarking via Mutual-Teacher Collaboration with Adaptive Feature Modulation

  • 设计双向交互机制,让嵌入器与提取器协同训练
  • 在真实和AI生成数据上均实现更高提取准确率
  • 适合需要高鲁棒性水印系统的开发者使用

现有深度图像水印方法采用固定的嵌入-失真-提取流程,嵌入器与提取器仅通过最终损失弱耦合,且独立优化。这种设计缺乏显式协作机制,导致嵌入器无法利用解码反馈,提取器也无法引导嵌入过程。为此,本文重新思考深度水印架构,将嵌入与提取重构为显式协作组件。提出协同交互机制(CIM),建立嵌入器与提取器之间的直接双向通信,实现互教师训练与联合优化。在此基础上,设计自适应特征调制模块(AFMM),通过解耦调制结构与强度,实现内容感知的特征调节,在嵌入时聚焦稳定图像特征,提取时抑制宿主干扰。双侧AFMM构成闭环协作,使嵌入行为与提取目标对齐。该架构级重构改变了水印鲁棒性的学习方式:不再依赖大量失真模拟,而是通过嵌入与提取间的协调表示学习自然产生鲁棒性。在真实世界与AI生成数据集上的实验表明,所提方法在保持高感知质量的同时,始终优于现有最优方法,展现出强鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Existing deep image watermarking methods follow a fixed embedding-distortion-extraction pipeline, where the embedder and extractor are weakly coupled through a final loss and optimized in isolation. This design lacks explicit collaboration, leaving no structured mechanism for the embedder to incorporate decoding-aware cues or for the extractor to guide embedding during training. To address this architectural limitation, we rethink deep image watermarking by reformulating embedding and extraction as explicitly collaborative components. To realize this reformulation, we introduce a Collaborative Interaction Mechanism (CIM) that establishes direct, bidirectional communication between the embedder and extractor, enabling a mutual-teacher training paradigm and coordinated optimization. Built upon this explicitly collaborative architecture, we further propose an Adaptive Feature Modulation Module (AFMM) to support effective interaction. AFMM enables content-aware feature regulation by decoupling modulation structure and strength, guiding watermark embedding toward stable image features while suppressing host interference during extraction. Under CIM, the AFMMs on both sides form a closed-loop collaboration that aligns embedding behavior with extraction objectives. This architecture-level redesign changes how robustness is learned in watermarking systems. Rather than relying on exhaustive distortion simulation, robustness emerges from coordinated representation learning between embedding and extraction. Experiments on real-world and AI-generated datasets demonstrate that the proposed method consistently outperforms state-of-the-art approaches in watermark extraction accuracy while maintaining high perceptual quality, showing strong robustness and generalization.

图像水印协同训练特征调制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。