用相似属性选负样本,提升多模态实体链接精度
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
- 基于杰卡德距离筛选相似属性的负样本,避免简单特征干扰
- 在多个基准数据集上显著提升链接准确率,最高增益达3.2%
- 结合上下文生成多视角图像,增强视觉表征鲁棒性
多模态实体链接(MEL)研究普遍采用对比学习作为主要目标,但以往方法常以整批样本为负样本,可能依赖易识别特征而忽略实体间关键差异。本文提出一种新方法JD-CCL(基于杰卡德距离的条件对比学习),利用元信息筛选具有相似属性的负样本,使链接任务更具挑战性且更稳健。此外,针对提及与实体间视觉模态的多样性问题,提出CVaCPT(上下文视觉辅助可控块变换)方法,通过融合多视角合成图像和上下文文本表示,对块表征进行缩放与偏移,增强视觉表达能力。在多个基准MEL数据集上的实验表明,该方法有效提升了链接性能。
原文摘要 · Abstract (English)
Previous research on multimodal entity linking (MEL) has primarily employed contrastive learning as the primary objective. However, using the rest of the batch as negative samples without careful consideration, these studies risk leveraging easy features and potentially overlook essential details that make entities unique. In this work, we propose JD-CCL (Jaccard Distance-based Conditional Contrastive Learning), a novel approach designed to enhance the ability to match multimodal entity linking models. JD-CCL leverages meta-information to select negative samples with similar attributes, making the linking task more challenging and robust. Additionally, to address the limitations caused by the variations within the visual modality among mentions and entities, we introduce a novel method, CVaCPT (Contextual Visual-aid Controllable Patch Transform). It enhances visual representations by incorporating multi-view synthetic images and contextual textual representations to scale and shift patch representations. Experimental results on benchmark MEL datasets demonstrate the strong effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。