arXiv:2501.03612eess.AScs.SD2025-01中稿 · Computer Speech an…被引 5

无需说话人嵌入,联合实现目标说话人提取与个人语音活动检测。

Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection

  • 用交叉注意力获取帧级特征替代传统说话人嵌入。
  • 在LibriMix和SparseLibriMix上性能领先,CALLHOME真实录音也表现优异。
  • 适合多说话人重叠场景,适用于实际语音分析任务。

确定‘谁在何时说了什么’在真实应用中仍具挑战性。通常,说话人聚类(SD)用于解决‘谁在何时说话’的问题,而目标说话人提取(TSE)或目标说话人语音识别(TSASR)则用于解决‘谁说了什么’的问题。尽管已有研究通过结合SD与TSE系统取得良好效果,但两者在输出一致性与场景适配性方面仍存在差异。为此,本文提出一种无需说话人嵌入的通用目标说话人提取与个人语音活动检测模型(USEF-TP),联合执行TSE与个人语音活动检测(PVAD)。该模型利用交叉注意力机制获得的帧级特征作为说话人相关特征,而非依赖传统说话人嵌入。同时,采用具有场景感知能力的差异化损失函数的多任务学习算法,确保在不同说话人重叠程度下的鲁棒性能。实验结果表明,所提模型在LibriMix与SparseLibriMix数据集上的TSE与PVAD任务中均表现更优;在CALLHOME真实录音数据集上也展现出竞争力。

原文摘要 · Abstract (English)

Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spoke when,' while Target Speaker Extraction (TSE) or Target Speaker Automatic Speech Recognition (TSASR) techniques are utilized to resolve the issue of 'who spoke what.' Although some works have achieved promising results by combining SD and TSE systems, inconsistencies remain between SD and TSE regarding both output inconsistency and scenario mismatch. To address these limitations, we propose a Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection (USEF-TP) model that jointly performs TSE and Personal Voice Activity Detection (PVAD). USEF-TP leverages frame-level features obtained through a cross-attention mechanism as speaker-related features instead of using speaker embeddings as in traditional approaches. Additionally, a multi-task learning algorithm with a scenario-aware differentiated loss function is applied to ensure robust performance across various levels of speaker overlap. The experimental results show that our proposed USEF-TP model achieves superior performance in TSE and PVAD tasks on the LibriMix and SparseLibriMix datasets. The results on the CALLHOME dataset demonstrate the competitive performance of our model on real recordings.

说话人分离语音活动检测多任务学习无嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。