让模型从立体声中推理声音源的移动轨迹与方向变化
Spatial Audio Question Answering and Reasoning on Dynamic Source Movements
- 用动态音频增强生成多样化运动模式数据
- 思考模式使模型推理准确率提升5.1%(单事件场景)
- 适合研究语音定位与多模态推理的开发者
空间音频理解旨在让机器解析复杂听觉场景,尤其关注随时间移动的声音源。本文研究基于运动推理的空间音频问答(Spatial AQA),要求模型直接从立体声中推断物体运动、位置及方向变化。首先,提出一种以运动为中心的空间音频增强框架,从孤立的单声道音频事件合成多样化的运动模式,实现可控且可扩展的训练数据生成。其次,提出端到端多模态微调方法,引入思考模式,使音频-语言模型在预测答案前生成显式的中间推理步骤。第三,研究查询条件源分离作为预处理阶段的影响,并比较三种推理方案:无掩码、音频定位模型(AGM)和真值掩码。结果表明,推理能放大源分离的收益,在单一事件提问时,思考模式带来+5.1%的显著提升。这些发现揭示了运动建模、推理能力与分离质量之间的相互作用,为推进空间音频理解提供了新见解。
原文摘要 · Abstract (English)
Spatial audio understanding aims to enable machines to interpret complex auditory scenes, particularly when sound sources move over time. In this work, we study Spatial Audio Question Answering (Spatial AQA) with a focus on movement reasoning, where a model must infer object motion, position, and directional changes directly from stereo audio. First, we introduce a movement-centric spatial audio augmentation framework that synthesizes diverse motion patterns from isolated mono audio events, enabling controlled and scalable training data generation. Second, we propose an end-to-end multimodal finetuning approach with a thinking mode, which allows audio-language models to produce explicit intermediate reasoning steps before predicting an answer. Third, we investigate the impact of query-conditioned source separation as a preprocessing stage and compare three inference regimes: no masking, an audio grounding model (AGM), and ground-truth masks. Our results show that reasoning amplifies the benefits of source separation, with thinking mode showing significant improvement of +5.1% when a single event is present in the question. These findings highlight the interplay between movement modeling, reasoning, and separation quality, offering new insights for advancing spatial audio understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。