arXiv:2601.12700eess.AScs.SD2026-01被引 1

用变分推断提升音频问答模型的准确率与可信度

Improving Audio Question Answering with Variational Inference

  • 引入IVON优化器,显式建模权重不确定性
  • 准确率提升,过拟合现象显著减少
  • 适合需要可靠置信度的应用场景

变分推断(VI)为模型参数后验分布估计提供了理论框架,可显式建模优化过程中的权重不确定性。通过捕捉这种不确定性,VI能提升预测可靠性,获得更校准的输出。本文研究了将先进的变分在线牛顿(IVON)优化器应用于多模态大语言模型在音频问答任务上的微调,验证了其在复杂多模态理解与推理中的优势。结果表明,采用VI不仅提升了预测准确性,还显著改善了模型校准性能,降低了模型过自信程度。这些改进对需要风险敏感决策的场景(如选择性预测)具有重要意义,其中可靠的置信度估计至关重要。

原文摘要 · Abstract (English)

Variational inference (VI) provides a principled framework for estimating posterior distributions over model parameters, enabling explicit modeling of weight uncertainty during optimization. By capturing this uncertainty, VI improves the reliability of predictions, yielding better calibrated outputs. In this work, we investigate the benefits of VI for challenging multimodal understanding and reasoning by applying the Improved Variational Online Newton (IVON), a recent VI optimizer, to fine-tuning a multimodal large language model on audio question answering tasks. Our results show that VI not only enhances predictive accuracy but also significantly improves calibration, reducing the model's overconfidence. These advances further support risk-sensitive applications such as selective prediction, where reliable confidence estimates are crucial.

音频问答变分推断多模态模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。