arXiv:2509.10143eess.AScs.CL2025-09中稿 · ITG Conference on …被引 1

分析会议转录中语音分离的泄漏问题,发现泄漏不影响最终效果。

Error Analysis in a Modular Meeting Transcription System

  • 改进时序敏感的泄漏分析框架,定位跨通道干扰区域。
  • 泄漏主要出现在主讲人独白区,但被语音活动检测忽略,影响小。
  • 先进聚类方法比简单能量检测提升33%性能,差距源于非说话段识别不足。

会议转录近年来发展迅速,但仍存在性能瓶颈。本文扩展了先前用于分析语音分离中泄漏问题的框架,引入对时序局部性的敏感性。结果显示,在仅主讲人发声的区域,存在显著的跨通道泄漏。然而,这些泄漏部分大多被语音活动检测(VAD)忽略,因此对最终性能影响较小。通过对比不同分割策略,发现先进的说话人聚类方法相比简单的能量阈值法,可将与理想分割的差距缩小三分之一。此外,我们揭示了剩余差异的主要成因。该系统在仅用LibriSpeech数据训练识别模块的前提下,在LibriCSS数据集上达到当前最优性能。

原文摘要 · Abstract (English)

Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previously proposed framework for analyzing leakage in speech separation with proper sensitivity to temporal locality. We show that there is significant leakage to the cross channel in areas where only the primary speaker is active. At the same time, the results demonstrate that this does not affect the final performance much as these leaked parts are largely ignored by the voice activity detection (VAD). Furthermore, different segmentations are compared showing that advanced diarization approaches are able to reduce the gap to oracle segmentation by a third compared to a simple energy-based VAD. We additionally reveal what factors contribute to the remaining difference. The results represent state-of-the-art performance on LibriCSS among systems that train the recognition module on LibriSpeech data only.

语音分离会议转录误差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。