arXiv:2602.13748cs.CLcs.CV2026-02中稿 · ACM ICMR 2026

提出新框架提升多模态事件抽取在数据少时的性能

RMPL: Relation-aware Multi-task Progressive Learning with Stage-wise Training for Multimedia Event Extraction

  • 分阶段训练,融合单模态与跨模态监督学习事件表示
  • 在M2E2上使用不同视觉语言模型均实现稳定提升
  • 适合低资源下多模态事件抽取研究者参考

多模态事件抽取(MEE)旨在从包含文本和图像的文档中识别事件及其论元,需在不同模态间对齐事件语义。当前进展受限于标注数据稀缺。尽管M2E2是唯一基准,但仅提供评估标注,无法直接用于监督训练。现有方法主要依赖跨模态对齐或基于VLM的推理时提示,未能显式学习结构化事件表示,且在多模态场景中论元定位较弱。为此,我们提出关系感知的多任务渐进学习框架RMPL,适用于低资源下的MEE。RMPL通过分阶段训练,整合来自单模态事件抽取和多模态关系抽取的异构监督信号。模型首先在统一模式下训练,学习跨模态共享的事件中心表示;随后利用混合文本与视觉数据微调事件提及识别与论元角色抽取。在M2E2基准上,使用多个VLM的实验表明,该方法在不同模态设置下均取得一致改进。

原文摘要 · Abstract (English)

Multimedia Event Extraction (MEE) aims to identify events and their arguments from documents that contain both text and images. It requires grounding event semantics across different modalities. Progress in MEE is limited by the lack of annotated training data. M2E2 is the only established benchmark, but it provides annotations only for evaluation. This makes direct supervised training impractical. Existing methods mainly rely on cross-modal alignment or inference-time prompting with Vision--Language Models (VLMs). These approaches do not explicitly learn structured event representations and often produce weak argument grounding in multimodal settings. To address these limitations, we propose RMPL, a Relation-aware Multi-task Progressive Learning framework for MEE under low-resource conditions. RMPL incorporates heterogeneous supervision from unimodal event extraction and multimedia relation extraction with stage-wise training. The model is first trained with a unified schema to learn shared event-centric representations across modalities. It is then fine-tuned for event mention identification and argument role extraction using mixed textual and visual data. Experiments on the M2E2 benchmark with multiple VLMs show consistent improvements across different modality settings.

事件抽取多模态低资源VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。