构建首个多模态大模型贝叶斯低秩适配基准,评估不确定性校准能力。
Bayesian Adaptation Gym: A Benchmark for the Bayesian Low-Rank Adaptation of Multi-Modal Language Models

- 采用贝叶斯低秩适配,实现大模型不确定性建模的高效计算
- 在多任务、多数据集上验证方法在分布外鲁棒性与主动学习中的表现
- 开源完整工具链,支持新方法对比与可复现研究
大型多模态语言模型正被广泛应用于高风险领域,校准良好的不确定性至关重要。传统贝叶斯方法需对所有模型参数进行后验近似,对现代大模型而言不可行。为此,近期研究转向贝叶斯低秩适配以实现可计算的后验近似。由于缺乏标准化评估基准,这些方法的实际收益尚不明确。为此,我们提出贝叶斯适配健身房(Bayesian Adaptation Gym, BAG),一个面向多模态语言模型贝叶斯适配的基准测试平台。BAG 提供经典贝叶斯基线与前沿适配方法的参考实现,配套多模态数据集与任务套件,用于探测模型校准度、分布偏移下的鲁棒性以及基于主动学习的决策能力。我们利用 BAG 在不同模型规模、数据集和任务上开展全面实验,系统分析当前贝叶斯适配方法的优势与局限。为推动后续研究,BAG 已完全开源:https://github.com/SRI-CSL/BayesAdapt。
原文摘要 · Abstract (English)
Large multi-modal language models are increasingly deployed in high-stakes domains, making well-calibrated uncertainty essential. Traditional Bayesian methods approximate posteriors over all model weights, which becomes intractable for modern large models. For this reason, recent work instead considers Bayesian low-rank adaptation to enable tractable posterior approximation. Due to a lack of a standardized benchmark to evaluate these approaches, it remains unclear where these methods provide meaningful benefits. To fill this gap, we introduce Bayesian Adaptation Gym (BAG), a benchmark for the Bayesian adaptation of multi-modal language models. BAG provides reference implementations of classic Bayesian baselines and state-of-the-art adaptation methods, along with a multi-modal dataset and task suite designed to probe calibration, robustness under distribution shift, and decision-making under uncertainty via active learning. Using BAG, we conduct and report extensive experiments across model sizes, datasets, and tasks to highlight the successes and failures of current Bayesian adaptation approaches. To enable further research, BAG is fully open source: https://github.com/SRI-CSL/BayesAdapt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。