arXiv:2504.00750cs.SDcs.LG2025-04中稿 · IEEE Journal of Se…被引 5

通过上下文与置信度感知提升语音提取质量

$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction

  • 引入跨模态上下文关联的掩码-恢复策略,实现全局推理
  • 在VoxCeleb2上多模型测试均显著提升提取效果
  • 针对低质量段落动态优化,适合复杂噪声场景

视听目标说话人分离(AV-TSE)旨在模拟人类利用视觉线索增强听觉感知的能力。尽管近期已提出多种模型,但多数仅依赖声学特征内的局部依赖关系,未能充分利用人类通过上下文推断模糊语音部分的能力。这一局限导致性能不理想且输出质量不一致,部分语句段落存在提取不佳或干扰说话人抑制不足的问题。为此,我们提出一种模型无关的策略——掩码-恢复(MAR),融合跨模态与模态内上下文相关性,实现提取模块内的全局推理。此外,为更好聚焦每个样本中的挑战性部分,我们设计细粒度置信度评分(FCS)模型,评估提取质量并引导模块重点改进低质量片段。在VoxCeleb2数据集上,采用六种主流AV-TSE主干网络验证所提训练范式,各项指标均实现一致提升。

原文摘要 · Abstract (English)

Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recently, most of them estimate target signals by primarily relying on local dependencies within acoustic features, underutilizing the human-like capacity to infer unclear parts of speech through contextual information. This limitation results in not only suboptimal performance but also inconsistent extraction quality across the utterance, with some segments exhibiting poor quality or inadequate suppression of interfering speakers. To close this gap, we propose a model-agnostic strategy called the Mask-And-Recover (MAR). It integrates both inter- and intra-modality contextual correlations to enable global inference within extraction modules. Additionally, to better target challenging parts within each sample, we introduce a Fine-grained Confidence Score (FCS) model to assess extraction quality and guide extraction modules to emphasize improvement on low-quality segments. To validate the effectiveness of our proposed model-agnostic training paradigm, six popular AV-TSE backbones were adopted for evaluation on the VoxCeleb2 dataset, demonstrating consistent performance improvements across various metrics.

语音分离多模态上下文建模置信度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。