低频语音未必安全,滤波缺失会严重泄露隐私。
Revisiting the Privacy of Low-Frequency Speech Signals: Exploring Resampling Methods, Evaluation Scenarios, and Speaker Characteristics
- 通过调整采样率与滤波器设计,测试语音隐私保护效果。
- 800 Hz以下采样仍可准确识别多数语音内容。
- 缺少抗混叠滤波会显著降低语音隐私,尤其对男性语音影响大。
现实中的音频记录虽能揭示社交行为,但也引发个人敏感数据的隐私担忧。本文研究将音频限制在低频以保护语音内容的有效性。针对不同采样率的音频重采样,比较了抗混叠滤波的影响。隐私保护效果通过自动语音识别(ASR)模型的词错误率(WER)衡量,实用性则由语音活动检测(VAD)模型评估。实验结果表明,在干净录音中,采样率高达800 Hz时,训练后的模型仍能正确转录大部分语音内容。同时分析了说话人性别与音高的影响,证实缺少抗混叠滤波会更严重损害语音隐私。
原文摘要 · Abstract (English)
While audio recordings in real life provide insights into social dynamics and conversational behavior, they also raise concerns about the privacy of personal, sensitive data. This article explores the effectiveness of restricting recordings to low-frequency audio to protect spoken content. For resampling the audio signals to different sampling rates, we compare the effect of employing anti-aliasing filtering. Privacy enhancement is measured by an increased word error rate of automatic speech recognition models. The impact on utility performance is measured with voice activity detection models. Our experimental results show that for clean recordings, models trained with a sampling rate of up to 800 Hz transcribe the majority of words correctly. For both models, we analyzed the impact of the speaker's sex and pitch, and we demonstrated that missing anti-aliasing filters more strongly compromise speech privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。