arXiv:2506.22789cs.SDcs.AI2025-06被引 3

用信息论方法让语音嵌入更公平隐私,同时保留有用内容

WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing

  • 通过互信息优化编码器,主动过滤说话人身份等敏感信息
  • 在3个数据集上降低敏感属性关联度81%,保留97%任务相关性
  • 适合关注语音隐私、公平性及自监督模型的开发者

语音嵌入常包含说话人身份、口音或人口统计信息等敏感属性,可能引发模型偏见和隐私泄露。本文提出WavShape,一种基于信息论的语音表示学习框架,在保障下游任务信息的同时,实现公平性与隐私保护。利用Donsker-Varadhan公式估计互信息,指导编码器系统性地剔除敏感属性,同时保留对任务关键的语音内容。在三个公开数据集上的实验表明,WavShape可将嵌入与敏感属性间的互信息降低高达81%,同时保留97%的任务相关信息。该工作融合信息论与自监督语音模型,推动了公平、隐私友好且资源高效的语音系统发展。

原文摘要 · Abstract (English)

Speech embeddings often retain sensitive attributes such as speaker identity, accent, or demographic information, posing risks in biased model training and privacy leakage. We propose WavShape, an information-theoretic speech representation learning framework that optimizes embeddings for fairness and privacy while preserving task-relevant information. We leverage mutual information (MI) estimation using the Donsker-Varadhan formulation to guide an MI-based encoder that systematically filters sensitive attributes while maintaining speech content essential for downstream tasks. Experimental results on three known datasets show that WavShape reduces MI between embeddings and sensitive attributes by up to 81% while retaining 97% of task-relevant information. By integrating information theory with self-supervised speech models, this work advances the development of fair, privacy-aware, and resource-efficient speech systems.

语音表示隐私保护信息论公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。