arXiv:2504.17353cs.CLcs.CV2025-04被引 1

首次将互增强效应拓展到多模态信息抽取,实现图文协同提升。

M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction

  • 提出多模态互增强新任务(M-MRE),构建首个支持该任务的数据集。
  • 设计适配多种大模型的提示格式适配器(PFA),有效提升跨模态理解性能。
  • 验证互增强在图文场景中同样有效,适合多模态理解与模型可解释性研究者。

互增强效应(MRE)是信息抽取与模型可解释性交叉领域的新兴方向,旨在通过不同粒度任务间的相互理解,联合建模以提升粗粒度与细粒度任务表现。尽管MRE已在文本领域得到探索和验证,其在视觉与多模态领域的适用性仍未知。本文首次将MRE扩展至多模态信息抽取领域,提出新的多模态互增强效应(M-MRE)任务,并构建相应数据集。为应对挑战,我们进一步提出提示格式适配器(PFA),完全兼容多种大视觉语言模型(LVLMs)。实验表明,在多模态文本-图像理解场景中,MRE现象依然存在,证明其可在三个相关任务间实现相互增益,证实了该效应超越文本域的普适性。

原文摘要 · Abstract (English)

Mutual Reinforcement Effect (MRE) is an emerging subfield at the intersection of information extraction and model interpretability. MRE aims to leverage the mutual understanding between tasks of different granularities, enhancing the performance of both coarse-grained and fine-grained tasks through joint modeling. While MRE has been explored and validated in the textual domain, its applicability to visual and multimodal domains remains unexplored. In this work, we extend MRE to the multimodal information extraction domain for the first time. Specifically, we introduce a new task: Multimodal Mutual Reinforcement Effect (M-MRE), and construct a corresponding dataset to support this task. To address the challenges posed by M-MRE, we further propose a Prompt Format Adapter (PFA) that is fully compatible with various Large Vision-Language Models (LVLMs). Experimental results demonstrate that MRE can also be observed in the M-MRE task, a multimodal text-image understanding scenario. This provides strong evidence that MRE facilitates mutual gains across three interrelated tasks, confirming its generalizability beyond the textual domain.

多模态互增强信息抽取大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。