arXiv:2608.04949cs.CVcs.CL2026-08中稿 · ACM MM2026

通过不确定性引导的模态增强与分布校准,提升多模态关系抽取的准确性和鲁棒性。

UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction

论文配图:UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction
图 1 · 摘自论文原文
  • 基于变分信息瓶颈建模特征为高斯分布,实现噪声过滤与语义保持。
  • 在三个基准数据集上达到最新性能,显著提升跨模态关系识别效果。
  • 模块可插拔,适用于需要处理模态异构和噪声干扰的多模态任务。

统一多模态关系抽取(UMRE)旨在识别文本实体与视觉对象之间的模态内与跨模态关系。现有研究面临两大关键问题:忽略固有随机不确定性导致噪声传播,以及不同模态分布间的深层异质性阻碍对齐。为此,本文提出不确定性引导的统一多模态关系抽取网络(UG-UMRE)。设计不确定性驱动的单模态增强(UDUA)模块,基于变分信息瓶颈将特征建模为高斯分布,并引入不确定性感知的自监督对比学习机制,在有效过滤噪声的同时保持语义一致性。进一步提出联合随机不确定性对齐(JAUA)模块,作为全局语义预校准机制,利用概率分布一致性构建共享潜在空间,通过同步跨模态统计特性消除分布差距,为细粒度交互奠定稳健基础。在三个基准数据集(UMRE、MORE、MNRE)上的实验表明,UG-UMRE取得当前最优性能。进一步分析验证了所提UDUA与JAUA模块的可插拔性与有效性。

原文摘要 · Abstract (English)

Unified Multimodal Relation Extraction (UMRE) aims to identify intra-modal and cross-modal relations between textual entities and visual objects. However, existing UMRE studies still encounter two critical issues: ignoring inherent aleatoric uncertainty causes noise propagation, and deep-seated heterogeneity between distinct modal distributions hinders alignment. To address these issues, we propose the Uncertainty-Guided UMRE Network (UG-UMRE). Specifically, we design an Uncertainty-Driven Unimodal Augmentation (UDUA) module, which models features as Gaussian distributions based on the Variational Information Bottleneck. By incorporating an uncertainty-aware self-supervised contrastive learning mechanism, UDUA effectively filters out noise while maintaining semantic consistency. Furthermore, we introduce the Joint Aleatoric Uncertainty Alignment (JAUA) module as a global semantic pre-calibration mechanism. JAUA leverages probabilistic distribution consistency to construct a shared latent space, eliminating the distributional gap by synchronizing cross-modal statistical properties, thereby laying a robust foundation for fine-grained interaction. Experiments on three benchmark datasets (UMRE, MORE, and MNRE) demonstrate that UG-UMRE achieves state-of-the-art performance. Further analysis validates the pluggable and effective performance of the proposed UDUA and JAUA modules.

多模态关系抽取不确定性建模分布对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。