arXiv:2508.15481cs.IR2025-08被引 1

首次系统评估多模态实体链接模型在视觉攻击下的鲁棒性,提出增强抗干扰能力的新方法。

On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking

  • 用大视觉模型提取初始实体描述,再通过网络检索动态生成候选描述。
  • 在五大数据集上提升准确率0.4%至35.7%,尤其在对抗攻击下表现显著。
  • 构建并开源首个MEL对抗样本数据集,适合关注多模态安全的研究者。

多模态数据的爆发式增长推动了多模态实体链接(MEL)模型的快速发展。然而,现有研究尚未系统考察视觉对抗攻击对MEL模型的影响。本文首次在多种对抗攻击场景下全面评估主流MEL模型的鲁棒性,涵盖图像到文本(I2T)和图像+文本到文本(IT2T)两个核心任务。实验结果表明,当前MEL模型普遍对视觉扰动缺乏足够鲁棒性。有趣的是,输入中的上下文语义信息可部分缓解对抗扰动的影响。基于此,我们提出一种基于大语言模型与检索增强的实体链接方法(LLM-RetLink),通过两阶段流程显著提升抗干扰能力:首先利用大视觉模型(LVMs)提取初始实体描述,再通过基于网络的检索动态生成候选描述句子。在五个数据集上的实验表明,LLM-RetLink将MEL准确率提升0.4%至35.7%,尤其在对抗条件下优势明显。本研究揭示了MEL鲁棒性的未被探索维度,构建并发布了首个MEL对抗样本数据集,为未来提升多模态系统在对抗环境中的韧性奠定了基础。

原文摘要 · Abstract (English)

The explosive growth of multimodal data has driven the rapid development of multimodal entity linking (MEL) models. However, existing studies have not systematically investigated the impact of visual adversarial attacks on MEL models. We conduct the first comprehensive evaluation of the robustness of mainstream MEL models under different adversarial attack scenarios, covering two core tasks: Image-to-Text (I2T) and Image+Text-to-Text (IT2T). Experimental results show that current MEL models generally lack sufficient robustness against visual perturbations. Interestingly, contextual semantic information in input can partially mitigate the impact of adversarial perturbations. Based on this insight, we propose an LLM and Retrieval-Augmented Entity Linking (LLM-RetLink), which significantly improves the model's anti-interference ability through a two-stage process: first, extracting initial entity descriptions using large vision models (LVMs), and then dynamically generating candidate descriptive sentences via web-based retrieval. Experiments on five datasets demonstrate that LLM-RetLink improves the accuracy of MEL by 0.4%-35.7%, especially showing significant advantages under adversarial conditions. This research highlights a previously unexplored facet of MEL robustness, constructs and releases the first MEL adversarial example dataset, and sets the stage for future work aimed at strengthening the resilience of multimodal systems in adversarial environments.

多模态对抗攻击实体链接鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。