arXiv:2601.05011cs.SDcs.LG2026-01被引 2

用预测熵自动调整提示权重,让零样本音频分类更稳定

Leveraging Prediction Entropy for Automatic Prompt Weighting in Zero-Shot Audio-Language Classification

  • 根据预测熵低的提示赋予更高权重,提升分类置信度
  • 在5个数据集上比传统提示集成方法准确率提升最高达5倍
  • 无需额外标注,计算开销极小,适合实际部署

音频-语言模型最近通过自然语言监督实现了强大的零样本音频事件分类能力,但其性能对文本提示的措辞极为敏感,微小变化会导致准确率大幅波动。已有工作通过提示学习或提示集成缓解此问题,但前者需标注数据,后者未考虑某些提示可能降低性能。本文提出一种基于熵引导的提示加权方法,通过最小化预测熵来生成新提示权重,以低熵作为高置信度的代理。该方法可应用于单个样本或批量音频,无需额外标签且计算开销极低。在涵盖环境音、城市音和人声的5个音频分类数据集上的实验表明,在零样本设置下,该方法相比经典提示集成方法取得一致提升,整体基准上准确率提升达5倍。

原文摘要 · Abstract (English)

Audio-language models have recently demonstrated strong zero-shot capabilities by leveraging natural-language supervision to classify audio events without labeled training data. Yet, their performance is highly sensitive to the wording of text prompts, with small variations leading to large fluctuations in accuracy. Prior work has mitigated this issue through prompt learning or prompt ensembling. However, these strategies either require annotated data or fail to account for the fact that some prompts may negatively impact performance. In this work, we present an entropy-guided prompt weighting approach that aims to find a robust combination of prompt contributions to maximize prediction confidence. To this end, we formulate a tailored objective function that minimizes prediction entropy to yield new prompt weights, utilizing low-entropy as a proxy for high confidence. Our approach can be applied to individual samples or a batch of audio samples, requiring no additional labels and incurring negligible computational overhead. Experiments on five audio classification datasets covering environmental, urban, and vocal sounds, demonstrate consistent gains compared to classical prompt ensembling methods in a zero-shot setting, with accuracy improvements 5-times larger across the whole benchmark.

零样本音频分类提示工程熵优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。