提出跨模态不确定性推理新任务与校准方法,实现多模态概率预测。
Unified Multimodal Uncertain Inference
- 构建跨文本、音频、视频的统一不确定性推理框架,支持多模态条件下的概率估计。
- 在多模态设置下,3B参数模型性能达到甚至超过32B参数基线模型。
- 引入校准机制(CLUE),提升模型输出概率的可信度,适合多模态认知研究者。
我们提出统一多模态不确定性推理(UMUI),一个涵盖文本、音频和视频的多模态推理任务,要求模型基于任意模态或其组合的前提,对假设生成校准的概率估计。尽管不确定性推理已在文本领域被探索,但扩展到其他模态仅限于单模态二元蕴含判断,缺乏跨模态细粒度概率推理的框架。为此,我们构建了一个由人工标注的评估集,包含音频、视觉及音视频场景下的标量概率判断,并在现有文本与音频基准上进行额外评估。我们提出CLUE(校准潜在不确定性估计),结合自洽教师校准与分布置信度探测,生成校准预测。实验表明,我们的3B参数模型在所有模态上的表现等同或优于高达32B参数的基线模型。
原文摘要 · Abstract (English)
We introduce Unified Multimodal Uncertain Inference (UMUI), a multimodal inference task spanning text, audio, and video, where models must produce calibrated probability estimates of hypotheses conditioned on a premise in any modality or combination. While uncertain inference has been explored in text, extension to other modalities has been limited to single-modality binary entailment judgments, leaving no framework for fine-grained probabilistic reasoning in or across other modalities. To address this, we curate a human-annotated evaluation set with scalar probability judgments across audio, visual, and audiovisual settings, and additionally evaluate on existing text and audio benchmarks. We introduce CLUE (Calibrated Latent Uncertainty Estimation), which combines self-consistent teacher calibration and distribution-based confidence probing to produce calibrated predictions. We demonstrate that our 3B-parameter model achieves equivalent or stronger performance than baselines up to 32B parameters across all modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。