arXiv:2501.03045eess.AScs.AI2025-01中稿 · ICASSP2025

针对室内外环境的单通道语音分离,提升移动端实时性与能效

Single-Channel Distance-Based Source Separation for Mobile GPU in Outdoor and Indoor Environments

  • 采用两阶段Conformer块与线性关系感知自注意力机制
  • 在移动GPU上实现更快推理速度与更高能效,支持实时处理
  • 特别优化室外复杂声学场景下的语音分离效果

本研究聚焦于室外环境中基于距离的语音分离(DSS)方法。与以往多集中于室内环境的研究不同,所提模型专为捕捉室外音频源的独特特性而设计。模型引入了两阶段Conformer块、线性关系感知自注意力(RSA)以及TensorFlow Lite GPU delegate。尽管线性RSA未能像二次型RSA那样显式建模物理线索,但其增强了模型对上下文的理解能力,从而在需要理解物理线索的室内外场景下表现更优。实验表明,该模型克服了现有方法的局限,显著提升了移动设备上的能效与实时推理速度。

原文摘要 · Abstract (English)

This study emphasizes the significance of exploring distance-based source separation (DSS) in outdoor environments. Unlike existing studies that primarily focus on indoor settings, the proposed model is designed to capture the unique characteristics of outdoor audio sources. It incorporates advanced techniques, including a two-stage conformer block, a linear relation-aware self-attention (RSA), and a TensorFlow Lite GPU delegate. While the linear RSA may not capture physical cues as explicitly as the quadratic RSA, the linear RSA enhances the model's context awareness, leading to improved performance on the DSS that requires an understanding of physical cues in outdoor and indoor environments. The experimental results demonstrated that the proposed model overcomes the limitations of existing approaches and considerably enhances energy efficiency and real-time inference speed on mobile devices.

语音分离移动端部署距离感知实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。