分离语音模型中的文本与声学特征,提升可解释性与隐私保护
Disentangling Textual and Acoustic Features of Neural Speech Representations
- 基于信息瓶颈原理,将语音表征拆分为文本与声学两部分
- 在情感识别与说话人识别任务中量化各层特征贡献度
- 可用于定位关键语音帧,适合需要透明性和隐私保护的场景
神经语音模型构建出高度纠缠的内部表征,以分布式方式编码多种特征(如基频、音量、句法类别或词义内容)。这种复杂性使得难以追踪表征对文本与声学信息的依赖程度,也难以抑制可能带来隐私风险(如性别或说话人身份)的声学特征编码。本文基于信息瓶颈原理,提出一种解耦框架,将复杂语音表征分解为两个独立成分:一个编码内容(即可转录为文本的部分),另一个编码与下游任务相关的声学特征。我们在情感识别和说话人识别任务上应用并评估该框架,量化了各模型层中文本与声学特征的贡献。此外,我们探索将该解耦框架作为归因方法,从文本与声学视角识别最具显著性的语音帧表征。
原文摘要 · Abstract (English)
Neural speech models build deeply entangled internal representations, which capture a variety of features (e.g., fundamental frequency, loudness, syntactic category, or semantic content of a word) in a distributed encoding. This complexity makes it difficult to track the extent to which such representations rely on textual and acoustic information, or to suppress the encoding of acoustic features that may pose privacy risks (e.g., gender or speaker identity) in critical, real-world applications. In this paper, we build upon the Information Bottleneck principle to propose a disentanglement framework that separates complex speech representations into two distinct components: one encoding content (i.e., what can be transcribed as text) and the other encoding acoustic features relevant to a given downstream task. We apply and evaluate our framework to emotion recognition and speaker identification downstream tasks, quantifying the contribution of textual and acoustic features at each model layer. Additionally, we explore the application of our disentanglement framework as an attribution method to identify the most salient speech frame representations from both the textual and acoustic perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。