arXiv:2509.12876cs.CLcs.MM2025-09中稿 · INLG 2025被引 4

评测大模型在多模态事件抽取中的表现并提出改进方法

Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents

  • 对比多个大模型在文本、图像和跨模态任务上的表现
  • 微调结合LoRA可显著提升模型性能,跨模态融合效果最优
  • 发现语义精度与跨模态对齐仍是主要挑战,适合多模态研究者参考

多媒体内容的爆发式增长亟需高效的多模态事件抽取(M2E2)系统。尽管大视觉语言模型(LVLMs)具备强大的跨模态能力,其在M2E2任务中的应用仍缺乏系统评估。本文首次对代表性LVLMs(包括DeepSeek-VL2和Qwen-VL系列)在M2E2数据集上进行了系统性评测,涵盖仅文本、仅图像和跨媒体子任务,在少样本提示与微调两种设置下进行评估。关键发现包括:(1)少样本LVLM在视觉任务上表现优异,但在文本任务上明显不足;(2)使用LoRA微调能显著提升模型性能;(3)多模态融合具有强协同效应,跨模态设置下表现最佳。我们还进行了详细错误分析,揭示了语义精度、定位准确性和跨模态对齐等持续存在的挑战,这些仍是推动M2E2发展的关键障碍。

原文摘要 · Abstract (English)

The proliferation of multimedia content necessitates the development of effective Multimedia Event Extraction (M2E2) systems. Though Large Vision-Language Models (LVLMs) have shown strong cross-modal capabilities, their utility in the M2E2 task remains underexplored. In this paper, we present the first systematic evaluation of representative LVLMs, including DeepSeek-VL2 and the Qwen-VL series, on the M2E2 dataset. Our evaluations cover text-only, image-only, and cross-media subtasks, assessed under both few-shot prompting and fine-tuning settings. Our key findings highlight the following valuable insights: (1) Few-shot LVLMs perform notably better on visual tasks but struggle significantly with textual tasks; (2) Fine-tuning LVLMs with LoRA substantially enhances model performance; and (3) LVLMs exhibit strong synergy when combining modalities, achieving superior performance in cross-modal settings. We further provide a detailed error analysis to reveal persistent challenges in areas such as semantic precision, localization, and cross-modal grounding, which remain critical obstacles for advancing M2E2 capabilities.

事件抽取多模态大模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。