分布式麦克风+几何信息,让声学场景理解更准更连贯。
Geometry-Informed Distributed Acoustic Scene Understanding

- 用分布麦克风和图神经网络融合声学时空特征
- 在模拟遮挡下空间一致性提升,优于集中式基线
- 适合需要物理一致场景推理的智能系统
多房间环境中的声学场景理解极具挑战。现有系统多依赖单一集中式麦克风阵列,常因墙体和门阻隔信号而失效。为此,我们提出一种几何感知的分布式声学场景理解框架。系统利用分布式麦克风,通过音频频谱图变换器与拓扑感知图神经网络融合时空声学特征,并解码为离散语义三元组。最后,冻结的大语言模型将这些符号化观测与环境几何信息结合,实现空间理解,推断合理缺失的事件转换,并生成物理上自洽的场景叙述。在自建多房间仿真器上的实验表明,该框架优于集中式基线,在模拟遮挡条件下显著提升空间一致性。
原文摘要 · Abstract (English)
Acoustic scene understanding in multi-room environments is a difficult task. Most existing systems use a single centralized microphone array, and they often fail because walls and doors block sound signals. To address this challenge, we propose a geometry-informed distributed acoustic scene understanding framework. Our system leverages distributed microphones and uses an audio spectrogram transformer and a topology-aware graph neural network to fuse spatio-temporal acoustic features. Then, these features are decoded into discrete semantic triplets. Finally, a frozen large language model combines these symbolic observations with the environmental geometry. This allows the system to perform spatial understanding, infer plausible missing transitions, and generate a physically consistent narrative of the scene. Experiments on a custom multi-room simulator demonstrate that our framework outperforms centralized baselines and improves spatial consistency under simulated occlusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。