用声音空间分布图增强视觉语言模型的场景理解能力
Acoustic Field Video for Multimodal Scene Understanding
- 将麦克风阵列采集的声音强度分布转为动态视频流
- 模型准确率从38.3%提升至67.4%,显著改善问答表现
- 适合做多模态感知、机器人导航与虚拟现实应用
我们提出一种新的多模态输入表征——声场视频。不同于传统视频(RGB+立体/单声道音频),声场视频通过空间化的方式可视化场景中的声音强度分布,为视觉语言模型提供全新的感知维度。我们的实时处理流程利用智能音箱、机器人及XR头显中常见的低成本波束成形麦克风阵列,但该传感能力在场景理解中尚未被充分利用。为评估空间声学信息的价值,我们构建了包含402个问答场景的评测集,对比了先进视觉语言模型在常规视频与配对声场视频下的表现。结果表明,引入空间声学数据后,模型准确率从38.3%提升至67.4%,且效果稳定一致。研究揭示:仅依赖视觉和音频输入时,许多日常场景理解任务仍存在信息不足的问题,而声场数据提供了可行且高效的新方向。视频演示见https://daehwakim.com/seeingsound
原文摘要 · Abstract (English)
We introduce and explore a new multimodal input representation for vision-language models: acoustic field video. Unlike conventional video (RGB with stereo/mono audio), our video stream provides a spatially grounded visualization of sound intensity across a scene, offering a new and powerful dimension of perceptual understanding. Our real-time pipeline uses low-cost beamforming microphone arrays that are already common in smart speakers and increasingly present in robotics and XR headsets, yet this sensing capability remains unutilized for scene understanding. To assess the value of spatial acoustic information, we constructed an evaluation set of 402 question-answer scenes, comparing a state-of-the-art VLM given conventional video with and without paired acoustic field video. Results show a clear and consistent improvement when incorporating spatial acoustic data; the VLM we test improves from 38.3% correct to 67.4%. Our findings highlight that many everyday scene understanding tasks remain underconstrained when relying solely on visual and audio input, and that acoustic field data provides a promising and practical direction for multimodal reasoning. A video demo is available at https://daehwakim.com/seeingsound
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。