arXiv:2604.16615cs.LGcs.AI2026-04

用音频上下文动态调节模型不确定性,提升语音相关任务的可靠性。

Beyond Feature Fusion: Contextual Bayesian PEFT for Multimodal Uncertainty Estimation

论文配图:Beyond Feature Fusion: Contextual Bayesian PEFT for Multimodal Uncertainty Estimation
图 1 · 摘自论文原文
  • 将音频特征作为上下文信号,动态调节低秩适配器的不确定性分布。
  • 在多种任务和主干网络上,性能优于纯文本微调与传统融合方法。
  • 适合需要高鲁棒性语音理解的低资源场景,如噪声环境下的语音识别。

我们提出 CoCo-LoRA,一种面向带音频上下文的文本预测任务的多模态、不确定性感知参数高效微调方法。现有参数高效微调(PEFT)方法如 LoRA 虽高效但通常为确定性,而近期基于贝叶斯的低秩适配器虽轻量建模不确定性,却仍主要依赖内部文本特征,难以反映背景噪声、信道差异或说话风格等外部声学因素带来的影响。CoCo-LoRA 通过在低秩空间中构建上下文变分后验,同时利用本地文本导出的适配器特征与音频导出的上下文信号进行联合建模。音频嵌入经一次投影进入共享上下文空间,再通过轻量级逐层头进行调制,实现从全局到局部、深度特定的不确定性调控,且无需高维多模态融合。随机性仅限于秩空间中的紧凑潜在成分,保持了参数效率,同时生成对音频敏感的异方差不确定性。在多种任务与主干结构上的评估表明,CoCo-LoRA 总体表现匹配或超越文本基线与传统特征融合方法,尤其在高覆盖率标签任务中优势显著。结果表明,将音频作为上下文不确定性信号而非融合特征流,是一种鲁棒且高效的多模态低资源预测方案。

原文摘要 · Abstract (English)

We introduce CoCo-LoRA, a multimodal, uncertainty-aware parameter-efficient fine-tuning method for text prediction tasks accompanied by audio context. Existing PEFT approaches such as LoRA are efficient but typically deterministic, while recent Bayesian low-rank adapters model uncertainty in a lightweight way yet remain largely unimodal and condition uncertainty primarily on internal text features. This leaves them poorly equipped to reflect uncertainty driven by external acoustic factors such as background noise, channel variability, or speaking style, which can materially affect reliability in speech-centered applications. CoCo-LoRA addresses this gap by conditioning a contextual variational posterior in the low-rank space on both local text-derived adapter features and an audio-derived context signal. A pooled audio embedding is projected once into a shared context space and then adapted through lightweight layer-wise heads, enabling global-to-local, depth-specific modulation of the adapter uncertainty and update without high-dimensional multimodal fusion. Stochasticity is confined to a compact latent component in the rank space, preserving PEFT scalability while producing audio-sensitive, heteroscedastic uncertainty. Based on our evaluations across diverse tasks and backbone combinations, CoCo-LoRA consistently matches or outperforms text-only PEFT and conventional feature-fusion transfer baselines, particularly on high-coverage labels where reliable adaptation is critical. The results indicate that using audio as a contextual uncertainty signal, rather than as a fused feature stream, provides a robust and parameter-efficient alternative for multimodal low-resource prediction.

多模态不确定性估计参数高效微调音频上下文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。