用大模型生成的声学场景训练轻量级DOA模型,提升实际应用泛化能力。
DOA Estimation with Lightweight Network on LLM-Aided Simulated Acoustic Scenes
- 基于大模型生成多样化声学场景,替代传统合成数据
- 轻量模型在多环境测试中保持高精度与低计算开销
- 适合嵌入式设备等资源受限场景部署
方向到达(DOA)估计在空间音频与声学信号处理中至关重要,广泛应用在真实世界中。现有大多数DOA模型依赖于将纯净语音与混响脉冲响应(RIRs)卷积生成的合成数据进行训练,因声学多样性有限而泛化能力受限。本文重新审视了基于大语言模型(LLMs)辅助生成的新型数据集上的DOA估计问题,该数据集提供了更真实、多样化的空间音频场景。我们在该数据集上基准测试了多种代表性神经网络DOA方法,并提出LightDOA,一种基于深度可分离卷积的轻量级模型,专为多通道输入和不同环境设计。实验表明,LightDOA在各类声学场景中均表现出良好的准确性和鲁棒性,同时保持极低的计算复杂度。本研究不仅展示了借助LLM生成的空间音频在推动鲁棒高效DOA估计方面的潜力,也验证了LightDOA作为资源受限场景下的高效解决方案的可行性。
原文摘要 · Abstract (English)
Direction-of-Arrival (DOA) estimation is critical in spatial audio and acoustic signal processing, with wide-ranging applications in real-world. Most existing DOA models are trained on synthetic data by convolving clean speech with room impulse responses (RIRs), which limits their generalizability due to constrained acoustic diversity. In this paper, we revisit DOA estimation using a recently introduced dataset constructed with the assistance of large language models (LLMs), which provides more realistic and diverse spatial audio scenes. We benchmark several representative neural-based DOA methods on this dataset and propose LightDOA, a lightweight DOA estimation model based on depthwise separable convolutions, specifically designed for mutil-channel input in varying environments. Experimental results show that LightDOA achieves satisfactory accuracy and robustness across various acoustic scenes while maintaining low computational complexity. This study not only highlights the potential of spatial audio synthesized with the assistance of LLMs in advancing robust and efficient DOA estimation research, but also highlights LightDOA as efficient solution for resource-constrained applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。