arXiv:2608.25289cs.SDcs.AI2026-08

用神经编码水印实现语音篡改检测与修复,支持无损恢复。

Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs

论文配图:Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs
图 1 · 摘自论文原文
  • 将水印嵌入神经编码表示中,实现内容篡改定位。
  • 四种篡改类型下均100%无误恢复原始内容。
  • 无需训练数据,适合低比特率语音系统使用。

语音片段局部篡改(如仅修改部分语句)给内容完整性验证带来挑战,因篡改比例越小,检测与定位越困难。水印提供主动防御方案:在分发前嵌入辅助信息。传统基于哈希的方案虽能精准检测定位,但一旦篡改则无法恢复原内容。本文基于自嵌入音频隐写框架,首次在理想条件下探索主动防御性能,从帧级定位、多比特最低有效位变体、及多种超低比特率神经编码器表征三方面展开研究。通过嵌入紧凑的神经编码表示而非加密哈希,该框架不仅可实现篡改区域的重建,还支持无需对抗样本的免训练检测与定位。在四种受控篡改类型、理想信道条件下实验表明,嵌入载荷即近似真实内容始终完全恢复且无比特错误。结果表明,神经编码器的选择是决定检测与定位性能的关键因素。

原文摘要 · Abstract (English)

Partial manipulation of speech recordings, where only localized segments of an utterance are altered, poses a significant challenge for content integrity verification, as reliable detection and localization of such edits becomes harder as the manipulated proportion decreases. Watermarking offers a proactive defense alternative by embedding auxiliary information prior to distribution; classical hash-based schemes achieve near-perfect detection and localization under ideal conditions, but the original content cannot be recovered once a segment is manipulated. Building on a prior self-embedding audio steganography framework, this work presents an initial exploration of proactive defense performance under ideal conditions, extending the investigation along three axes: frame-level localization, multi-bit least significant bit variants, and evaluation across multiple ultra-low-bitrate neural codec representations. By embedding a compact neural codec representation rather than a cryptographic hash, the framework additionally enables recovery of the manipulated regions, while supporting training-free detection and localization without spoofed examples. Experiments across four controlled manipulation types under ideal channel conditions show that the embedded payload, and hence an approximate reconstruction of the authentic content, is always fully recovered without bit errors. The results also indicate that the choice of neural codec is the dominant factor for detection and localization performance.

音频水印神经编码内容验证篡改修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。