arXiv:2512.07845cs.SDcs.AI2025-12

构建含空间信息的音频3D场景数据集,助力音频引导的空间理解研究。

AudioScene: Integrating Object-Event Audio into 3D Scenes

  • 用大模型辅助+人工验证,高效生成带空间标注的音频数据。
  • 在音频定位与零样本机器人导航任务中表现优于现有方法。
  • 适合研究音频感知、智能体导航与多模态学习的学者。

音频分析的快速发展展现出其在人机交互、环境监测和公共安全中的巨大潜力,但现有音频数据集普遍缺乏空间上下文。为填补这一空白,我们构建了两个新颖的音视频空间场景数据集:AudioScanNet 和 AudioRoboTHOR,旨在探索3D环境中音频条件下的各类任务。通过将音频片段与空间对齐的3D场景结合,我们的数据集支持研究音频信号如何与空间信息相互作用。为将音频事件与空间位置关联,我们利用大语言模型的常识推理能力,并辅以严格的人工验证,该方法相比纯人工标注更具可扩展性,同时保持高精度、完整性和多样性,通过标注者间一致性及两项基准任务——基于音频的3D视觉定位与基于音频的机器人零样本导航——进行了量化评估。结果揭示了当前音频中心方法的局限性,凸显了本数据集在推动音频引导的空间学习中的实际挑战与重要意义。

原文摘要 · Abstract (English)

The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two novel audiospatial scene datasets, AudioScanNet and AudioRoboTHOR, designed to explore audioconditioned tasks within 3D environments. By integrating audio clips with spatially aligned 3D scenes, our datasets enable research on how audio signals interact with spatial context. To associate audio events with corresponding spatial information, we leverage the common sense reasoning ability of large language models and supplement them with rigorous human verification, This approach offers greater scalability compared to purely manual annotation while maintaining high standards of accuracy, completeness, and diversity, quantified through inter annotator agreement and performance on two benchmark tasks audio based 3D visual grounding and audio based robotic zeroshot navigation. The results highlight the limitations of current audiocentric methods and underscore the practical challenges and significance of our datasets in advancing audio guided spatial learning.

音频理解3D场景多模态机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。