无需标签,通过自博弈提升音频细粒度推理能力
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

- 构建无标签音频对比游戏,模型自动生成线索并识别差异
- 在TREA等数据集上显著提升事件顺序、重复等细粒度推理性能
- 适合需要高精度音频理解的场景,如智能听觉系统
大型音频语言模型(LALMs)在声学理解方面进展迅速,但仍难以处理细粒度音频推理任务(如识别事件顺序、重复和持续时间)。现有后训练方法严重依赖昂贵的外部标注或仅提供粗粒度语义信号。为此,我们提出Audio-Zero,首个面向LALMs的无标签自演化框架,可提升细粒度听觉感知与推理能力。Audio-Zero从无标签音频对比对中构建听觉自博弈游戏:多数参与者听到参考音频,而一个异常听众听到细微变体。模型首先生成描述所听内容的线索,再通过推理不一致线索来识别异常者。由于异常者由构造决定,游戏可提供可验证奖励,无需人工标注答案。在Qwen2-Audio-7B-Instruct与Qwen2.5-Omni-7B模型上,于TREA、MMAU Test-mini和MMAR数据集上的实验表明,Audio-Zero在保持广泛音频理解能力的同时,显著提升了细粒度音频推理性能。演化与诊断分析进一步揭示,在游戏压力下,模型自然涌现出越来越精细的听觉描述。
原文摘要 · Abstract (English)
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-Zero constructs an auditory self-play game from unlabeled audio contrast pairs: most players hear a reference audio, while one odd listener hears a subtle variant. The model first generates clues describing what it hears and then identifies the odd listener by reasoning over inconsistencies among clues. Since the odd listener is known by construction, the game provides verifiable rewards without any annotated answers. Experiments with Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA, MMAU Test-mini and MMAR show that Audio-Zero improves fine-grained audio reasoning while preserving broad audio understanding. Evolutionary and diagnostic analyses further reveal that increasingly fine-grained auditory descriptions emerge naturally from game pressure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。