arXiv:2508.04425eess.AScs.SD2025-08被引 13

通过分离说话人与文本特征,实现语音验证的文本自适应

Text adaptation for speaker verification with speaker-text factorized embeddings

  • 将语音分解为说话人和文本嵌入,再融合生成定制化表示
  • 仅需少量无特定文本的语音样本,即可显著提升文本不匹配场景下的识别率
  • 适合需要快速适配新语料的实用语音验证系统

训练或注册数据与实际测试数据之间的文本不匹配会严重损害文本依赖型语音验证系统的性能。尽管可通过精心采集目标语料来解决该问题,但成本高且灵活性差。本文提出一种新型文本自适应框架:构建说话人-文本因子化网络,将输入语音分解为说话人嵌入和文本嵌入,并在后期阶段融合为统一表征。利用少量与说话人无关的适配语音,可提取目标文本嵌入,将原本与文本无关的说话人嵌入转换为针对特定文本的嵌入。RSR2015实验表明,该方法在文本不匹配条件下能显著提升系统性能。

原文摘要 · Abstract (English)

Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system performance. Although this problem can be solved by carefully collecting data with the target speech content, such data collection could be costly and inflexible. In this paper, we propose a novel text adaptation framework to address the text mismatch issue. Here, a speaker-text factorization network is proposed to factorize the input speech into speaker embeddings and text embeddings and then integrate them into a single representation in the later stage. Given a small amount of speaker-independent adaptation utterances, text embeddings of target speech content can be extracted and used to adapt the text-independent speaker embeddings to text-customized speaker embeddings. Experiments on RSR2015 show that text adaptation can significantly improve the performance of text mismatch conditions.

语音验证文本自适应嵌入分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。