评测11种模型在声音来源识别任务中的表现,揭示大模型的推理盲区。
Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification

- 按任务输入方式分四类评估,避免不同方法直接比拼
- 最佳模型类别准确率85.6%,细粒度仅56.7%
- 大模型虽自信但常答错,且对小样本不敏感
我们评测了11种音频分类方法:五种面向任务的闭集LLM(四个Gemini模型和开源的Kimi-Audio-7B-Instruct)、四种固定词汇标签器(YAMNet、PANNs、Whisper-AT、SSLAM)、一个零样本音文模型(CLAP)以及一个音频引导的LLM(BAT)。在包含2,242段音频、23个细粒度类别和11个大类别的闭集声音来源识别任务上进行评估。由于这些方法在任务接收方式和输出评分机制上存在根本差异,我们将其分为四类评估层级,报告每层级的宏平均精确率、召回率、F1和误报率。最佳模型Gemini-3.1-Pro-Preview在类别层面达到85.6% F1,细粒度为56.7%。Kimi-Audio在同等规模下表现良好,类别层面达67.5% F1,细粒度为32.9%,但无法回答1.6%的样本。SSLAM与CLAP在未见候选列表情况下达到或超过最优闭集模型的类别层面性能,但在细粒度层面落后。分析Gemini模型在8,968次响应中的思维链发现,回应长度不能预测准确性;看似“整体判断优于详细分析”的现象实为难度混淆所致;错误答案被自信陈述的比例高达92%至100%。我们提供所有11种方法的完整类别混淆矩阵和指标,揭示精度损失的主要结构性错误模式,并为选择方法家族提供实用建议。
原文摘要 · Abstract (English)
We benchmark eleven audio classification methods: five task-aware closed-set LLMs (four Gemini models plus open-weight Kimi-Audio-7B-Instruct), four fixed-vocabulary taggers (YAMNet, PANNs, Whisper-AT, and SSLAM), a zero-shot audio-text model (CLAP), and an audio-grounded LLM (BAT). We evaluate them on a closed-set sound-source identification task over 2,242 clips spanning 23 fine-grained classes and 11 categories. Since these methods differ fundamentally in how they receive the task and how outputs are scored, we group them into four evaluation tiers rather than one leaderboard, reporting macro Precision, Recall, F1, and false-negative rate per tier. The best model, Gemini-3.1-Pro-Preview, reaches 85.6 percent category-level F1 and 56.7 percent fine-grained F1. Kimi-Audio is competitive for its size, reaching 67.5 percent category-level F1 and 32.9 percent fine-grained F1, but fails to answer 1.6 percent of samples. SSLAM and CLAP match or exceed the best closed-set model at the category level without seeing the candidate list, but fall behind at the fine-grained level. Analyzing the Gemini models' chain-of-thought across 8,968 responses, we find that response length does not predict accuracy, an apparent "holistic judgment beats detailed analysis" effect is better explained as a difficulty confound, and wrong answers are stated confidently 92 to 100 percent of the time. We report full per-class confusion matrices and metrics for all eleven methods, identify the structural error modes behind most of the accuracy loss between granularities, and give practical guidance for choosing among these method families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。