arXiv:2602.20592cs.SDeess.AS2026-02中稿 · Interspeech 2026

用信息论方法量化语音中多维度独立性,揭示情感与语言病理特征的编码差异。

Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning

  • 引入带边界约束的神经互信息估计,结合非参数验证量化语音特征间依赖关系。
  • 跨维度互信息低于0.15纳特,源-滤波器间互信息达0.47纳特,显示弱耦合性。
  • 情感由声源主导(80%),语言与病理由滤波器主导(60%和58%),适合表征学习研究者。

语音信号在共享声学通道中编码情绪、语言及病理信息;然而,解耦性通常通过下游任务性能间接评估。本文提出一种信息论框架,通过结合有界神经互信息(MI)估计与非参数验证,量化手工特征中多维度间的统计依赖性。在六个语料库上,跨维度互信息保持较低水平(<0.15纳特),估计边界紧密,表明数据中耦合较弱,而源-滤波器间互信息显著更高(0.47纳特)。归因分析显示,情绪维度中源成分贡献80%的总互信息,语言与病理维度分别由滤波器成分主导(60%和58%)。该研究为语音中维度独立性的量化提供了原则性方法。

原文摘要 · Abstract (English)

Speech signals encode emotional, linguistic, and pathological information within a shared acoustic channel; however, disentanglement is typically assessed indirectly through downstream task performance. We introduce an information-theoretic framework to quantify cross-dimension statistical dependence in handcrafted acoustic features by integrating bounded neural mutual information (MI) estimation with non-parametric validation. Across six corpora, cross-dimension MI remains low, with tight estimation bounds ($< 0.15$ nats), indicating weak statistical coupling in the data considered, whereas Source--Filter MI is substantially higher (0.47 nats). Attribution analysis, defined as the proportion of total MI attributable to source versus filter components, reveals source dominance for emotional dimensions (80\%) and filter dominance for linguistic and pathological dimensions (60\% and 58\%, respectively). These findings provide a principled framework for quantifying dimensional independence in speech.

语音解耦信息论表征学习互信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。