arXiv:2603.26766cs.CV2026-03

用视觉不可察觉的水印技术,让屏幕拍摄后仍能准确提取信息。

JND-Guided Neural Watermarking with Spatial Transformer Decoding for Screen-Capture Robustness

  • 通过模拟真实屏幕拍摄失真,训练网络对抗复杂干扰。
  • 结合人眼感知阈值,将水印藏在视觉不敏感区域以保画质。
  • 自动修复画面扭曲和裁剪,适合实际应用部署。

屏幕拍摄鲁棒水印旨在将可提取信息不可察觉地嵌入宿主图像,使水印在屏幕显示与相机重拍的复杂失真流程中仍能保留。然而,在保持良好视觉质量的同时实现高提取准确率仍是难题,主要因屏幕拍摄通道引入严重且交织的退化,包括摩尔纹、色域偏移、透视畸变和传感器噪声。本文提出端到端深度学习框架,联合优化水印嵌入与提取。创新点包括:(i) 全面的噪声模拟层,真实建模屏幕拍摄失真,特别是基于物理的摩尔纹生成器,通过对抗训练使网络学习全谱失真下的鲁棒表征;(ii) 基于人眼可察觉失真(JND)的感知损失函数,通过监督JND系数图与水印残差间的感知差异,自适应调节嵌入强度,将水印能量集中于人眼不敏感区域以最大化视觉质量;(iii) 两个互补的自动定位模块——基于语义分割的前景提取用于重拍图像校正,对称噪声模板机制用于抗裁剪区域恢复,支持在真实场景下全自动解码。大量实验表明,该方法在水印图像上平均达到30.94 dB PSNR 和0.94 SSIM,同时嵌入127比特载荷。

原文摘要 · Abstract (English)

Screen-shooting robust watermarking aims to imperceptibly embed extractable information into host images such that the watermark survives the complex distortion pipeline of screen display and camera recapture. However, achieving high extraction accuracy while maintaining satisfactory visual quality remains an open challenge, primarily because the screen-shooting channel introduces severe and entangled degradations including Moiré patterns, color-gamut shifts, perspective warping, and sensor noise. In this paper, we present an end-to-end deep learning framework that jointly optimizes watermark embedding and extraction for screen-shooting robustness. Our framework incorporates three key innovations: (i) a comprehensive noise simulation layer that faithfully models realistic screen-shooting distortions -- notably including a physically-motivated Moiré pattern generator -- enabling the network to learn robust representations against the full spectrum of capture-channel noise through adversarial training; (ii) a Just Noticeable Distortion (JND) perceptual loss function that adaptively modulates watermark embedding strength by supervising the perceptual discrepancy between the JND coefficient map and the watermark residual, thereby concentrating watermark energy in perceptually insensitive regions to maximize visual quality; and (iii) two complementary automatic localization modules -- a semantic-segmentation-based foreground extractor for captured image rectification and a symmetric noise template mechanism for anti-cropping region recovery -- that enable fully automated watermark decoding under realistic deployment conditions. Extensive experiments demonstrate that our method achieves an average PSNR of 30.94~dB and SSIM of 0.94 on watermarked images while embedding 127-bit payloads.

水印视觉感知屏幕拍摄

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。