arXiv:2412.18024cs.LG2024-12被引 9

提出可处理多模态冲突的不确定性量化方法,提升关键场景决策可靠性。

Multimodal Learning with Uncertainty Quantification based on Discounted Belief Fusion

  • 采用基于冲突的折扣机制,动态重分配不可靠模态的不确定质量。
  • 在高冲突场景下显著优于现有方法,准确识别矛盾样本。
  • 方法对模态顺序不变,适用于多源信息融合任务。

多模态人工智能模型在医疗、金融和自动驾驶等领域广泛应用,其信息来自图像、文本、音频、视频等多种来源。然而,有效管理由噪声、证据不足或模态间冲突引发的不确定性,对可靠决策至关重要。现有不确定性感知学习方法(如证据平均)在高冲突场景下低估不确定性,且当前最先进的证据平均策略不具备顺序不变性,难以扩展到多模态场景。为此,我们提出一种新型多模态学习方法,具备顺序不变的证据融合能力,并引入基于冲突的折扣机制,在检测到不可靠模态时重新分配不确定质量。通过理论分析与实验验证,该方法能有效区分冲突与非冲突样本,相较于以往模型,在基于不确定性的冲突检测任务中表现更优。

原文摘要 · Abstract (English)

Multimodal AI models are increasingly used in fields like healthcare, finance, and autonomous driving, where information is drawn from multiple sources or modalities such as images, texts, audios, videos. However, effectively managing uncertainty - arising from noise, insufficient evidence, or conflicts between modalities - is crucial for reliable decision-making. Current uncertainty-aware machine learning methods leveraging, for example, evidence averaging, or evidence accumulation underestimate uncertainties in high-conflict scenarios. Moreover, the state-of-the-art evidence averaging strategy is not order invariant and fails to scale to multiple modalities. To address these challenges, we propose a novel multimodal learning method with order-invariant evidence fusion and introduce a conflict-based discounting mechanism that reallocates uncertain mass when unreliable modalities are detected. We provide both theoretical analysis and experimental validation, demonstrating that unlike the previous work, the proposed approach effectively distinguishes between conflicting and non-conflicting samples based on the provided uncertainty estimates, and outperforms the previous models in uncertainty-based conflict detection.

多模态学习不确定性量化冲突检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。