系统梳理端到端多说话人语音识别新进展,助力复杂对话场景理解。
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
- 提出SIMO与SISO两种架构分类,分析其优劣与适用场景。
- 对比主流方法在标准数据集上的性能,揭示关键改进方向。
- 适合语音识别、智能会议等需要区分多人对话的研究者参考。
单声道多说话人自动语音识别(ASR)因数据稀缺及重叠语音中词句归属难题而面临挑战。近年来,端到端(E2E)架构取代级联系统,减少错误传播并更好地融合语音内容与说话人身份信息。尽管进展迅速,该领域仍缺乏全面综述。本文系统梳理了E2E多说话人ASR的神经网络方法,提出分类体系,分析:(1) 预分割音频下的SIMO与SISO架构范式及其权衡;(2) 基于两类范式的最新架构与算法改进;(3) 长时语音处理中的分段策略与说话人一致假设拼接技术。进一步在标准基准上评估并比较各类方法。最后讨论开放挑战与未来方向,推动构建更鲁棒、可扩展的多说话人ASR系统。
原文摘要 · Abstract (English)
Monaural multi-speaker automatic speech recognition (ASR) remains challenging due to data scarcity and the intrinsic difficulty of recognizing and attributing words to individual speakers, particularly in overlapping speech. Recent advances have driven the shift from cascade systems to end-to-end (E2E) architectures, which reduce error propagation and better exploit the synergy between speech content and speaker identity. Despite rapid progress in E2E multi-speaker ASR, the field lacks a comprehensive review of recent developments. This survey provides a systematic taxonomy of E2E neural approaches for multi-speaker ASR, highlighting recent advances and comparative analysis. Specifically, we analyze: (1) architectural paradigms (SIMO vs.~SISO) for pre-segmented audio, analyzing their distinct characteristics and trade-offs; (2) recent architectural and algorithmic improvements based on these two paradigms; (3) extensions to long-form speech, including segmentation strategy and speaker-consistent hypothesis stitching. Further, we (4) evaluate and compare methods across standard benchmarks. We conclude with a discussion of open challenges and future research directions towards building robust and scalable multi-speaker ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。