用深度学习建模室内声学,提升声音场重建效果。
Deep, data-driven modeling of room acoustics: literature review and research perspectives
- 结合几何与波动信息构建声学模型,保留时空结构
- 在声音场重建任务中取得良好效果
- 适合声学、语音处理与智能空间研究者
日常听觉体验受室内环境声学特性影响。声学建模旨在建立声波在这些环境中传播的数学表示,对回声辅助听觉导航、嘈杂环境下语音理解恢复等应用具有重要意义。近年来,深度学习推动了多个学科范式转变,声学研究亦不例外。多数现有深度数据驱动模型借鉴语音与图像处理方法,缺乏声波传播的内在时空结构。近期,融合几何或波动信息的深度学习模型在声场重建任务中表现优异。本文系统综述深度数据驱动声学建模研究进展,将其置于与传统物理及数据驱动模型对比的框架中,并分析现有方法的优势与不足,指出未来关键挑战。
原文摘要 · Abstract (English)
Our everyday auditory experience is shaped by the acoustics of the indoor environments in which we live. Room acoustics modeling is aimed at establishing mathematical representations of acoustic wave propagation in such environments. These representations are relevant to a variety of problems ranging from echo-aided auditory indoor navigation to restoring speech understanding in cocktail party scenarios. Many disciplines in science and engineering have recently witnessed a paradigm shift powered by deep learning (DL), and room acoustics research is no exception. The majority of deep, data-driven room acoustics models are inspired by DL-based speech and image processing, and hence lack the intrinsic space-time structure of acoustic wave propagation. More recently, DL-based models for room acoustics that include either geometric or wave-based information have delivered promising results, primarily for the problem of sound field reconstruction. In this review paper, we will provide an extensive and structured literature review on deep, data-driven modeling in room acoustics. Moreover, we position these models in a framework that allows for a conceptual comparison with traditional physical and data-driven models. Finally, we identify strengths and shortcomings of deep, data-driven room acoustics models and outline the main challenges for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。