arXiv:2506.12032cs.LGcs.AI2025-06

用神经水印保护科学数据,抗压缩变形还能精准验证来源

Embedding Trust at Scale: Physics-Aware Neural Watermarking for Secure and Verifiable Data Pipelines

  • 用卷积自编码器将二进制信息嵌入温度、涡度等结构化数据中
  • 在噪声、裁剪、压缩下仍保持98%以上解码准确率,重建误差低于1%
  • 适合气候模拟、流体仿真等高维科学数据的可信溯源与审计

我们提出一种面向科学数据完整性的鲁棒神经水印框架,适用于气候建模和流体模拟等高维领域。通过卷积自编码器,将二进制消息隐式嵌入温度、涡度、大地形势等结构化数据中。该方法在噪声注入、裁剪和压缩等有损变换下仍能保持水印持久性,同时重建质量接近原始数据(均方误差低于1%)。相比传统奇异值分解(SVD)水印方法,在ERA5和纳维-斯托克斯数据集上实现超过98%的比特准确率,且视觉无差异。该系统为高性能科学工作流中的数据出处、可审计性和可追溯性提供了可扩展、模型兼容的工具,推动通过可验证的物理感知水印保障人工智能系统安全。评估基于具有物理基础的科学数据集,作为典型压力测试;该框架可自然扩展至卫星图像、自动驾驶感知流等其他结构化领域。

原文摘要 · Abstract (English)

We present a robust neural watermarking framework for scientific data integrity, targeting high-dimensional fields common in climate modeling and fluid simulations. Using a convolutional autoencoder, binary messages are invisibly embedded into structured data such as temperature, vorticity, and geopotential. Our method ensures watermark persistence under lossy transformations - including noise injection, cropping, and compression - while maintaining near-original fidelity (sub-1\% MSE). Compared to classical singular value decomposition (SVD)-based watermarking, our approach achieves $>$98\% bit accuracy and visually indistinguishable reconstructions across ERA5 and Navier-Stokes datasets. This system offers a scalable, model-compatible tool for data provenance, auditability, and traceability in high-performance scientific workflows, and contributes to the broader goal of securing AI systems through verifiable, physics-aware watermarking. We evaluate on physically grounded scientific datasets as a representative stress-test; the framework extends naturally to other structured domains such as satellite imagery and autonomous-vehicle perception streams.

神经水印数据溯源科学计算物理感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。