用听觉重建空间,让机器人靠声音感知环境。
BinauralVAE: Spatial Audio Reconstruction For World Models

- 基于变分自编码器构建双耳音频重建模型
- 在模拟环境中实现听觉与动作的因果关联建模
- 适合研究多模态世界模型或听觉导航的开发者
具身人工智能长期依赖视觉感知,导致大量以视觉为中心的世界模型涌现。然而,这种依赖无法完整捕捉空间理解,在视觉遮挡、低光或断电等场景下会暴露缺陷,此时声学信息成为关键的空间感知手段。尽管潜力巨大,但真实空间音频及以声音为核心的的世界模型研究仍十分有限。本文介绍 BinauralVAE:一个灵活开源的音频重建管道(https://github.com/Luizerko/BinauralVAE),探索从基础基线到数学严谨架构的多种模型。通过评估多种变分自编码器(包括复值变体),学习双耳信号的鲁棒潜在表示。该方法结合 AudioWorldSim,利用模拟机器人在环境中移动时采集的真实声学数据,建立未来听觉主导世界模型的状态表征基础,旨在映射导航动作与其声学后果之间的直接因果关系,推动声音作为空间认知的核心互补模态。
原文摘要 · Abstract (English)
Embodied artificial intelligence has historically very much relied on visual perception, leading to a proliferation of multiple vision-centric world models. However, this reliance fails to capture spatial understanding in its entirety and can even present vulnerabilities in environments with visual occlusions, low-light conditions, or blackouts-scenarios, where acoustic information becomes a critical alternative for spatial awareness and navigation. Despite its potential, research into realistic spatial audio and particularly the development of audio-centric world models remains sparse. In this technical report, we introduce BinauralVAE: a flexible, open-source pipeline (https://github.com/Luizerko/BinauralVAE) that explores multiple models for spatialized audio reconstruction, progressing from fundamental baselines to advanced, mathematically grounded architectures. Our approach evaluates various Variational Autoencoder architectures -- including complex-valued variants -- to learn robust latent representations of binaural signals. Developed alongside AudioWorldSim, our methodology leverages realistic acoustic data captured as a simulated robot navigates an environment. This pipeline establishes a foundation for state representation in a future audio-based world model, designed to map the direct causal connection between navigational actions and their resulting acoustic consequences, and helping to enable sound as an essential complementary modality for spatial knowledge acquisition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。