不依赖语音分离,用声音方向直接提取对话中每个人的说话内容。
Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR

- 用声源方向作为空间先验,指导语音分离
- 在真实和回放对话中提升语音识别准确率
- 无需语音分离步骤,适合长段多人对话场景
在长时多说话人对话中,说话人活动极不平衡且频繁重叠,导致难以确定“谁在何时说了什么”。滑动窗口连续语音分离(CSS)缓解了监督信号稀疏问题,但常出现跨窗口说话人不一致和残留串扰,实践中仍需语音分离来可靠分配说话人。受会议中说话人声源方向(DOA)稳定性的启发,我们提出PATSE,一种多通道位置感知目标说话人提取前端,利用DOA作为空间先验,直接提取每个目标说话人的语音。PATSE结合基于DOA的时空编码器与调节模块生成带说话人标识的语音流,通过简单后处理(如语音活动检测)即可推断说话人活动,无需显式语音分离。在回放和真实对话上的实验表明,其性能持续优于CSS及基于语音分离的流水线。
原文摘要 · Abstract (English)
In long-form multi-party conversations, highly imbalanced speaker activity and frequent overlap make it difficult to identify "who spoke when and what". Sliding-window continuous speech separation (CSS) mitigates sparse supervision, but often suffers from cross-window speaker inconsistency and residual crosstalk, which in practice requires diarization for reliable speaker attribution. Motivated by the stability of speakers' directions of arrival (DOAs) in meetings, we propose PATSE, a multi-channel Position-Aware Target Speaker Extraction front-end that uses DOA as a spatial prior to directly extract the speech of each target speaker. PATSE combines a DOA-guided spatial encoder and conditioner to generate speaker-attributed streams, from which speaker activity can be inferred via simple post-processing (e.g., VAD) without explicit diarization. Experiments on both replayed and real conversations show consistent ASR gains outperforming CSS and diarization-based pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。