用问答方式理解空间音频,让模型能回答声音事件的位置与关系。
Towards Spatial Audio Understanding via Question Answering
- 构建基于规则和大模型重述的语音问答数据集
- 仅用场景级标注,性能接近全帧级监督模型
- 适合研究音频场景理解与语言引导建模的学者
本文提出一种通过问答范式实现一阶混响场(FOA)信号空间音频理解的新框架,旨在将声事件定位与检测(SELD)拓展至空间场景理解和推理。我们基于规则方法为STARSS23数据集构建并发布细粒度时空文本描述,并利用大语言模型(LLM)提升语言多样性;同时构建了与STARSS23场景对齐的问答数据集,涵盖事件存在性、定位、空间与时间关系等维度;并通过LLM生成多版本问题重述以增强语言变体。我们还开发了一个基线空间音频问答模型,输入为FOA信号和自然语言问题,输出关于声事件发生、时空关系的分类答案。尽管仅使用场景级问答监督训练,其性能仍可媲美采用帧级时空标注的全监督SELD模型。结果表明,语言引导方法在空间音频理解中具有潜力,为融合语言监督到空间场景分析开辟新方向。
原文摘要 · Abstract (English)
In this paper, we introduce a novel framework for spatial audio understanding of first-order ambisonic (FOA) signals through a question answering (QA) paradigm, aiming to extend the scope of sound event localization and detection (SELD) towards spatial scene understanding and reasoning. First, we curate and release fine-grained spatio-temporal textual descriptions for the STARSS23 dataset using a rule-based approach, and further enhance linguistic diversity using large language model (LLM)-based rephrasing. We also introduce a QA dataset aligned with the STARSS23 scenes, covering various aspects such as event presence, localization, spatial, and temporal relationships. To increase language variety, we again leverage LLMs to generate multiple rephrasings per question. Finally, we develop a baseline spatial audio QA model that takes FOA signals and natural language questions as input and provides answers regarding various occurrences, temporal, and spatial relationships of sound events in the scene formulated as a classification task. Despite being trained solely with scene-level question answering supervision, our model achieves performance that is comparable to a fully supervised sound event localization and detection model trained with frame-level spatiotemporal annotations. The results highlight the potential of language-guided approaches for spatial audio understanding and open new directions for integrating linguistic supervision into spatial scene analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。