系统梳理空间音频技术进展,覆盖生成、理解与评估方法。
ASAudio: A Survey of Advanced Spatial Audio Research
- 按输入输出形式与任务类型分类整理空间音频研究
- 综述主流数据集、评价指标与基准测试体系
- 适合关注沉浸式音效与多模态交互的研究者
随着空间音频技术的快速发展,其在增强现实(AR)、虚拟现实(VR)等场景中受到广泛关注。相较于传统单声道音频,空间音频能提供更真实、沉浸的听觉体验。尽管该领域已取得显著进展,但缺乏对现有方法及核心技术的系统性综述。本文对空间音频研究进行全面回顾,按时间线梳理相关工作,并基于输入输出表示形式以及生成与理解任务进行分类,总结了空间音频研究的多个方面。此外,还系统介绍了相关数据集、评估指标与基准测试,从训练与评估双重视角提供深入见解。相关资源见 https://github.com/dieKarotte/ASAudio。
原文摘要 · Abstract (English)
With the rapid development of spatial audio technologies today, applications in AR, VR, and other scenarios have garnered extensive attention. Unlike traditional mono sound, spatial audio offers a more realistic and immersive auditory experience. Despite notable progress in the field, there remains a lack of comprehensive surveys that systematically organize and analyze these methods and their underlying technologies. In this paper, we provide a comprehensive overview of spatial audio and systematically review recent literature in the area. To address this, we chronologically outlining existing work related to spatial audio and categorize these studies based on input-output representations, as well as generation and understanding tasks, thereby summarizing various research aspects of spatial audio. In addition, we review related datasets, evaluation metrics, and benchmarks, offering insights from both training and evaluation perspectives. Related materials are available at https://github.com/dieKarotte/ASAudio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。