arXiv:2606.21305cs.SDcs.CL2026-06中稿 · Interspeech 2026

让语音嵌入可听懂,用无标签分解揭示说话人特征。

LISE : Listenable Interpretable Speaker Embeddings

论文配图:LISE : Listenable Interpretable Speaker Embeddings
图 1 · 摘自论文原文
  • 无标签分解预训练语音嵌入,生成可结构化表示。
  • 人类听辨准确率达83.9%,验证了可听解释性。
  • 保持原有验证性能,适合需要透明性的语音系统。

基于深度神经网络的自动说话人验证(ASV)系统表现优异,但其嵌入表示仍不透明,缺乏对编码语音特征的结构化与可感知解释。现有方法要么依赖说话人属性标注,要么引入未经听觉验证的替代表示。本文提出可听可解释的说话人嵌入(LISE),一种无标签框架,将预训练说话人嵌入分解为少量组件。该分解生成结构化表示,支持分析嵌入所编码的信息。LISE在x-vector和ECAPA-TDNN上保持近似原性能,等错误率(EER)下降可忽略。关键的是,听觉实验表明,参与者以83.9%准确率区分说话人,证实了这些组件对人类听觉的可解释性。

原文摘要 · Abstract (English)

Deep neural network-based automatic speaker verification (ASV) systems achieve impressive performance but their embedding representations remain opaque, lacking a structured and perceptually verifiable explanation of the vocal characteristics they encode. Existing approaches either require annotation of speaker attributes or introduce alternative representations whose interpretability is unvalidated with listeners. We propose Listenable Interpretable Speaker Embeddings (LISE), a label-free framework that decomposes pretrained speaker embeddings into a small set of components. This decomposition yields a structured representation that supports the analysis of what information has been encoded by speaker embeddings. LISE preserves ASV performance with negligible EER degradation on x-vector and ECAPA-TDNN. Crucially, the interpretability of these components for human listeners is demonstrated through listening experiments, where participants distinguished speakers with 83.9% accuracy.

语音验证可解释性嵌入分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。