分离语音内容与说话人特征,让语音识别更可解释。
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation
- 用多时序分辨率分解表示,分离语音内容和说话人特征
- 在语音识别准确率提升的同时,明确区分说话人身份
- 适合需要模型可解释性的语音分析场景
端到端的基于Transformer的自动语音识别(ASR)系统常在其学习表征中混杂多种语音特征,导致可解释性差。本研究提出可解释的解耦Transformer,通过不同时间分辨率将内部表示分解为显式的语音内容与说话人特征子嵌入。实验结果表明,所提出的解耦Transformer在实现语音识别性能提升的同时,能清晰分离说话人身份,适用于说话人聚类任务。
原文摘要 · Abstract (English)
End-to-end transformer-based automatic speech recognition (ASR) systems often capture multiple speech traits in their learned representations that are highly entangled, leading to a lack of interpretability. In this study, we propose the explainable Disentangled-Transformer, which disentangles the internal representations into sub-embeddings with explicit content and speaker traits based on varying temporal resolutions. Experimental results show that the proposed Disentangled-Transformer produces a clear speaker identity, separated from the speech content, for speaker diarization while improving ASR performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。