让音频模型像人一样理解声音细节,提升语义推理能力
AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound
- 基于人类认知构建听觉语义推理框架,引导模型深入理解声音细节
- 在多场景训练下超越现有最佳模型,显著提升细粒度音频语义理解性能
- 专为语义推理设计新数据集,解决零样本评估中的数据污染问题
音频语言模型在多种声音理解任务中表现优异,但在细粒度声音语义推理方面仍受限。本文提出 AudSemThinker,其推理机制基于受人类认知启发的听觉语义框架。为此,我们构建了 AudSem——一个专为音频-语言模型语义描述符推理而设计的新数据集。该数据集通过多阶段稳健流程生成标注,有效避免零样本评估中的数据污染问题。实验表明,AudSemThinker 在多种训练设置下均优于当前最优模型,展现出强大的语义音频推理能力。相关模型与数据集已公开发布。
原文摘要 · Abstract (English)
Audio-language models have shown promising results in various sound understanding tasks, yet they remain limited in their ability to reason over the fine-grained semantics of sound. In this paper, we present AudSemThinker, a model whose reasoning is structured around a framework of auditory semantics inspired by human cognition. To support this, we introduce AudSem, a novel dataset specifically curated for semantic descriptor reasoning in audio-language models. AudSem addresses the persistent challenge of data contamination in zero-shot evaluations by providing a carefully filtered collection of audio samples paired with captions generated through a robust multi-stage pipeline. Our experiments demonstrate that AudSemThinker outperforms state-of-the-art models across multiple training settings, highlighting its strength in semantic audio reasoning. Both AudSemThinker and the AudSem dataset are released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。