构建首个大规模多模态空间音频数据集,支持真实场景下声音定位与生成研究。
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations
- 采集四类真实场景的双耳/全向音频与多视角视频,同步标注语音、歌词等细粒度信息。
- 在5个基础任务中验证了数据集可实现高质量空间建模,如空间语音合成与声源定位准确率提升。
- 适合从事虚拟现实、空间音频生成与多模态理解的研究者使用。
人类依赖多感官整合感知三维空间环境,听觉线索对声音源定位至关重要。尽管空间音频在虚拟现实(VR/AR)等沉浸式技术中扮演关键角色,现有大多数多模态数据集仅提供单声道音频,限制了空间音频生成与理解的发展。为此,我们提出MRSAudio,一个大规模多模态空间音频数据集,旨在推动空间音频理解与生成研究。该数据集涵盖四个不同类别:MRSLife、MRSSpeech、MRSMusic 和 MRSSing,覆盖多样化的现实场景。数据包含同步的双耳与全向音频、外视角与内视角视频、运动轨迹,以及细粒度标注,如转录文本、音素边界、歌词、乐谱和提示词。为展示其价值,我们定义了五个基础任务:音频空间化、空间文本到语音、空间歌唱合成、空间音乐生成与声事件定位检测。结果表明,MRSAudio能实现高质量的空间建模,并支持广泛的空间音频研究。演示与数据获取地址:https://mrsaudio.github.io。
原文摘要 · Abstract (English)
Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR, most existing multimodal datasets provide only monaural audio, which limits the development of spatial audio generation and understanding. To address these challenges, we introduce MRSAudio, a large-scale multimodal spatial audio dataset designed to advance research in spatial audio understanding and generation. MRSAudio spans four distinct components: MRSLife, MRSSpeech, MRSMusic, and MRSSing, covering diverse real-world scenarios. The dataset includes synchronized binaural and ambisonic audio, exocentric and egocentric video, motion trajectories, and fine-grained annotations such as transcripts, phoneme boundaries, lyrics, scores, and prompts. To demonstrate the utility and versatility of MRSAudio, we establish five foundational tasks: audio spatialization, and spatial text to speech, spatial singing voice synthesis, spatial music generation and sound event localization and detection. Results show that MRSAudio enables high-quality spatial modeling and supports a broad range of spatial audio research. Demos and dataset access are available at https://mrsaudio.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。