arXiv:2510.09505eess.AS2025-10

融合声音方向信息的会议语音分段框架,提升多人对话分离效果。

Spatially-Augmented Sequence-to-Sequence Neural Diarization for Meetings

  • 将声源方向估计融入序列到序列模型,增强空间感知能力。
  • 离线模式下相对错误率降低7.4%,结合通道注意力提升超19%。
  • 适合需要高精度多人语音分离的会议记录与实时转写场景。

本文提出一种空间增强型序列到序列神经说话人分离(SA-S2SND)框架,将由SRP-DNN估计的到达方向(DOA)线索整合至S2SND主干网络中。采用两阶段训练策略:先用单通道音频和DOA特征训练模型,再在多通道输入下以DOA为指导进一步优化。此外,引入模拟DOA生成方案,缓解对匹配多通道语料的依赖。在AliMeeting数据集上,SA-S2SND持续优于S2SND基线,在离线模式下相对错误率(DER)降低7.4%,结合通道注意力时改善超过19%。结果表明,空间线索与跨通道建模高度互补,能在在线与离线设置中均取得优异性能。

原文摘要 · Abstract (English)

This paper proposes a Spatially-Augmented Sequence-to-Sequence Neural Diarization (SA-S2SND) framework, which integrates direction-of-arrival (DOA) cues estimated by SRP-DNN into the S2SND backbone. A two-stage training strategy is adopted: the model is first trained with single-channel audio and DOA features, and then further optimized with multi-channel inputs under DOA guidance. In addition, a simulated DOA generation scheme is introduced to alleviate dependence on matched multi-channel corpora. On the AliMeeting dataset, SA-S2SND consistently outperform the S2SND baseline, achieving a 7.4% relative DER reduction in the offline mode and over 19% improvement when combined with channel attention. These results demonstrate that spatial cues are highly complementary to cross-channel modeling, yielding good performance in both online and offline settings.

语音分离会议记录空间线索深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。