构建可生成真实双耳音频数据集的开源平台,助力音频世界模型研究。
AudioWorldSim: Realistic Binaural Audio Datasets For World Models

- 基于SoundSpaces 2.0扩展,自动执行随机智能体导航以生成音频数据
- 修复连续声音合成问题,提升音频真实性与一致性
- 开源平台支持复现,适合音频感知与世界模型研究者使用
本技术报告介绍AudioWorldSim,一个开源平台,旨在生成逼真的双耳音频数据集,推动基于音频的机器学习研究,尤其是世界模型的发展。该平台作为Meta SoundSpaces 2.0的定制扩展,利用其完整的声学框架,专注于自动执行随机智能体导航,并修复了连续声音合成中的关键问题。AudioWorldSim已公开发布于https://github.com/Luizerko/AudioWorldSim,供研究社区使用,以促进实验复现。
原文摘要 · Abstract (English)
This technical report presents AudioWorldSim, an open-source platform designed to generate realistic binaural audio datasets and advance research in audio-based machine learning, particularly world models. Built as a custom extension of Meta's SoundSpaces 2.0 platform, AudioWorldSim leverages their comprehensive acoustics framework, but focuses on the automatic rollout of random agent navigations, as well as implements crucial fixes to how continuous sound is composed. AudioWorldSim is made publicly available to the research community at https://github.com/Luizerko/AudioWorldSim to facilitate reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。