arXiv:2604.25591eess.AScs.AI2026-04被引 3

首次系统评估音频大模型的不确定性估计,发现语义级方法更可靠。

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

论文配图:Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models
图 1 · 摘自论文原文
  • 对比五种不确定性估计方法,聚焦语义与验证类策略
  • 语义级方法在音频推理任务中表现更优,尤其在复杂场景
  • 适用于需高可信度的场景,如幻觉检测与不可答问题识别

近期的音频感知大语言模型(ALLMs)在多样化的音频理解与推理任务中表现出强大能力,但仍频繁产生幻觉或过度自信的输出。尽管不确定性估计在纯文本大模型中已有广泛研究,但在涉及音频条件生成的ALLMs中仍缺乏系统探索,因感知模糊性和跨模态对齐引入了额外挑战。本文首次对ALLMs中的不确定性估计进行系统性实证研究,对比了五种代表性方法:预测熵、长度归一化熵、语义熵、离散语义熵和P(True),覆盖多个模型及多样化评估设置,涵盖通用音频理解、推理、幻觉检测和不可答问题问答。结果揭示两大关键发现:第一,在通用音频推理基准上,语义级与验证类方法持续优于词元级基线;第二,在可信度导向的基准上,不确定性方法的有效性显著依赖于模型与任务,表明通用推理结论无法直接迁移至幻觉与不可答场景。我们进一步探讨了基于不确定性的自适应推理作为潜在下游应用。本研究为构建可靠、具备不确定性感知的音视频系统提供了基础。

原文摘要 · Abstract (English)

Recent audio-aware large language models (ALLMs) have demonstrated strong capabilities across diverse audio understanding and reasoning tasks, but they still frequently produce hallucinated or overly confident outputs. While uncertainty estimation has been extensively studied in text-only LLMs, it remains largely unexplored for ALLMs, where audio-conditioned generation introduces additional challenges such as perceptual ambiguity and cross-modal grounding. In this work, we present the first systematic empirical study of uncertainty estimation in ALLMs. We benchmark five representative methods, including predictive entropy, length-normalized entropy, semantic entropy, discrete semantic entropy, and P(True), across multiple models and diverse evaluation settings spanning general audio understanding, reasoning, hallucination detection, and unanswerable question answering. Our results reveal two key findings. First, semantic-level and verification-based methods consistently outperform token-level baselines on general audio reasoning benchmarks. Second, on trustworthiness-oriented benchmarks, the relative effectiveness of uncertainty methods becomes notably more model- and benchmark-dependent, indicating that conclusions drawn from general reasoning settings do not straightforwardly transfer to hallucination and unanswerable-question scenarios. We further explore uncertainty-based adaptive inference as a potential downstream application. We hope this study provides a foundation for future research on reliable, uncertainty-aware audio-language systems.

音频理解不确定性估计大模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。