统一生成框架实现跨领域事件抽取,支持端到端与流水线两种模式。
A Scalable Cross-Domain Event Extraction System via a Unified Generative Training Framework

- 采用统一生成式序列到序列框架,联合完成事件检测与论元抽取。
- 单模型在多领域数据上微调,可同时保留领域特征并泛化至新标签空间。
- 提供可视化网页平台,支持文档上传、配置对比与跨域结果分析。
事件抽取是信息抽取的基础任务。以往方法常将事件检测与论元抽取分离,或依赖特定数据集设计,限制了可扩展性与跨领域泛化能力。我们提出一种统一的生成式序列到序列框架,可联合执行事件抽取子任务,并支持流水线与端到端两种配置。通过在多个跨领域事件数据集上微调预训练语言模型,使单一模型既能保持领域特异性语义,又能泛化至大规模且不断演化的标签空间。我们通过一个面向研究者与实践者的网络应用展示了这些能力:平台支持文档上传、基于模式的事件抽取、触发词与论元的可视化,以及跨领域的不同抽取配置比较。
原文摘要 · Abstract (English)
Event extraction is fundamental to information extraction. Prior approaches often separate event detection and argument extraction or depend on dataset-specific designs, limiting scalability and cross-domain generalization. We propose a unified generative sequence-to-sequence framework that performs event extraction subtasks jointly and supports both pipeline and end-to-end configurations. We fine-tune pretrained language models on multiple event datasets across diverse domains, enabling a single model to retain domain-specific semantics while generalizing over large and evolving label spaces. We demonstrate these capabilities through a web-based application tailored for researchers and practitioners. The platform supports document upload, schema-aware event extraction, visualization of triggers and arguments, and comparison of different extraction configurations across domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。