arXiv:2603.16805cs.SD2026-03

让音频分轨水印可恢复,通过联合训练提升分离后信息提取能力

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training

  • 分轨水印先分离再解码,用共享结构+独立密钥嵌入信息
  • 联合优化水印与分离器,后分离比特恢复率显著提升
  • 适用于音乐/语音混合场景,兼顾听觉质量与水印鲁棒性

现代音频常由多个音轨混合而成,提出在分离前对各音轨独立嵌入水印,并在分离后恢复所有水印的框架。研究发现,仅使用通用鲁棒水印与现成分离器的简单流程,会导致解码失败,说明对一般失真鲁棒并不等同于对分离伪影鲁棒。为此,构建一个受控验证流程,将分离器作为检测器的一部分,可与水印系统共同选择或优化。在语音+音乐和人声+伴奏混合数据集上的实验表明,该方法在保持良好听觉感知质量的同时,显著提升了分离后的水印恢复性能。

原文摘要 · Abstract (English)

Modern audio is created by mixing stems from different sources, raising the question: can we independently watermark each stem and recover all watermarks after separation? We study a separation-first, multi-stream watermarking framework --embedding distinct information into stems using unique keys but a shared structure, mixing, separating, and decoding from each output. A naive pipeline (robust watermarking + off-the-shelf separation) yields poor bit recovery, showing robustness to generic distortions does not ensure robustness to separation artifacts. To enable this, we study separation-aware watermarking in a controlled verification pipeline, where the separator is part of the detector and can be selected or optimized together with the watermarking system. Experiments on speech+music and vocal+accompaniment mixtures show substantial gains in post-separation recovery while maintaining perceptual quality.

音频水印多流处理联合训练分离鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。