arXiv:2507.08412cs.SDeess.AS2025-07被引 1

通过反转波形段保护录音中的语音隐私,同时保留环境音质。

Enforcing Speech Content Privacy in Environmental Sound Recordings using Segment-wise Waveform Reversal

  • 用语音活动检测与分离定位语音段,逐段反转以破坏可懂性
  • 语音可懂度降低97.9% WER,音源识别准确率仅降2.7% SCAD
  • 保留高感知质量(FAD=1.40),适合需隐私保护的音频数据共享

环境声音录音常包含可识别语音,引发隐私担忧,限制数据的分析、共享与再利用。本文提出一种方法,在保持声景完整性和整体音频质量的前提下,使语音不可理解。该方法通过反转波形段来扭曲语音内容,并结合语音活动检测与语音分离流程,实现更精准的语音定位。为验证效果,采用三部分评估协议:1)使用词错误率(WER)评估语音可懂度;2)使用预训练模型的声源分类准确率下降(SCAD)评估音源可检测性;3)使用弗雷歇音频距离(FAD)评估音频质量,基于包含未篡改语音的参考数据集。在由语音与环境音线性混合构成的模拟数据集上实验表明,该方法实现97.9%的高语音可懂度损失(WER),仅2.7%的音源识别性能下降(SCAD),且感知质量优异(FAD=1.40)。消融实验验证了各模块贡献。此外,引入随机拼接可增强抗恢复鲁棒性,代价为轻微音质下降。

原文摘要 · Abstract (English)

Environmental sound recordings often contain intelligible speech, raising privacy concerns that limit analysis, sharing and reuse of data. In this paper, we introduce a method that renders speech unintelligible while preserving both the integrity of the acoustic scene, and the overall audio quality. Our approach involves reversing waveform segments to distort speech content. This process is enhanced through a voice activity detection and speech separation pipeline, which allows for more precise targeting of speech. In order to demonstrate the effectivness of the proposed approach, we consider a three-part evaluation protocol that assesses: 1) speech intelligibility using Word Error Rate (WER), 2) sound sources detectability using Sound source Classification Accuracy-Drop (SCAD) from a widely used pre-trained model, and 3) audio quality using the Fréchet Audio Distance (FAD), computed with our reference dataset that contains unaltered speech. Experiments on this simulated evaluation dataset, which consists of linear mixtures of speech and environmental sound scenes, show that our method achieves satisfactory speech intelligibility reduction (97.9% WER), minimal degradation of the sound sources detectability (2.7% SCAD), and high perceptual quality (FAD of 1.40). An ablation study further highlights the contribution of each component of the pipeline. We also show that incorporating random splicing to our speech content privacy enforcement method can enhance the algorithm's robustness to attempt to recover the clean speech, at a slight cost of audio quality.

语音隐私音频安全声学场景波形处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。