arXiv:2501.05729cs.SDcs.AI2025-01中稿 · IEEE Signal Proces…被引 8

让语音验证结果可解释,像人工鉴定一样分析发音特征。

ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification

  • 引入发音特征分析,模拟人工语音比对过程。
  • 可生成语音嵌入并可视化具体发音特点。
  • 适合需要透明度的司法或安全场景使用。

在语音验证中,我们采用计算方法判断一段语音是否属于注册说话人。该任务类似于人工法医语音比对,需结合语言学分析与听觉测量来对比和评估语音样本。尽管已有诸多进展,但尚未构建出能提供类似人工比对解释性的语音验证系统。本文提出一种新型可解释发音特征导向(ExPO)网络,引入描述说话人在音素层面特征的发音特质,类比法医语音比对的做法。ExPO不仅能生成语句级说话人嵌入,还能实现发音特质的细粒度分析与可视化,提供可解释的语音验证流程。此外,我们从说话人内与说话人间变异角度研究发音特质,确定最有效的验证特征,为可解释语音验证迈出重要一步。代码已开源:https://github.com/mmmmayi/ExPO。

原文摘要 · Abstract (English)

In speaker verification, we use computational method to verify if an utterance matches the identity of an enrolled speaker. This task is similar to the manual task of forensic voice comparison, where linguistic analysis is combined with auditory measurements to compare and evaluate voice samples. Despite much success, we have yet to develop a speaker verification system that offers explainable results comparable to those from manual forensic voice comparison. A novel approach, Explainable Phonetic Trait-Oriented (ExPO) network, is proposed in this paper to introduce the speaker's phonetic trait which describes the speaker's characteristics at the phonetic level, resembling what forensic comparison does. ExPO not only generates utterance-level speaker embeddings but also allows for fine-grained analysis and visualization of phonetic traits, offering an explainable speaker verification process. Furthermore, we investigate phonetic traits from within-speaker and between-speaker variation perspectives to determine which trait is most effective for speaker verification, marking an important step towards explainable speaker verification. Our code is available at https://github.com/mmmmayi/ExPO.

语音验证可解释性发音特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。