融合图文信息的事件抽取方法,提升开放域事件识别能力
A Multimodal Text- and Graph-Based Approach for Open-Domain Event Extraction from Documents

- 结合大模型文本表示与图结构建模文档级上下文
- 在多个数据集上超越现有最先进方法,支持未见事件类型
- 适合需要跨领域事件理解的应急决策与文档分析场景
事件抽取对事件理解与分析至关重要,支持文档摘要和应急决策等任务。现有方法存在局限:(1) 封闭域算法仅限预定义事件类型,难以泛化至新类型;(2) 开放域方法虽能处理无约束事件类型,却未充分挖掘大语言模型(LLMs)的能力,且未显式建模文档级上下文、结构与语义推理,导致长文本中出现注意力稀释和信息丢失问题。为此,我们提出多模态开放域事件抽取方法 MODEE,通过结合基于图的学习与大模型文本表示,实现文档级推理建模。在大规模数据集上的实证评估表明,MODEE 在开放域事件抽取上优于现有最先进方法,并可泛化至封闭域任务,性能超过已有算法。
原文摘要 · Abstract (English)
Event extraction is essential for event understanding and analysis. It supports tasks such as document summarization and decision-making in emergency scenarios. However, existing event extraction approaches have limitations: (1) closed-domain algorithms are restricted to predefined event types and thus rarely generalize to unseen types and (2) open-domain event extraction algorithms, capable of handling unconstrained event types, have largely overlooked the potential of large language models (LLMs) despite their advanced abilities. Additionally, they do not explicitly model document-level contextual, structural, and semantic reasoning, which are crucial for effective event extraction but remain challenging for LLMs due to lost-in-the-middle phenomenon and attention dilution. To address these limitations, we propose multimodal open-domain event extraction, MODEE , a novel approach for open-domain event extraction that combines graph-based learning with text-based representation from LLMs to model document-level reasoning. Empirical evaluations on large datasets demonstrate that MODEE outperforms state-of-the-art open-domain event extraction approaches and can be generalized to closed-domain event extraction, where it outperforms existing algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。