用编码器-解码器模型精准定位文本中事件的时间与地点。
When and Where Did it Happen? An Encoder-Decoder Model to Identify Scenario Context

- 基于流行病学论文数据集,微调编码器-解码器模型识别事件时空信息。
- 小规模微调模型在时间地点预测上优于通用大模型和语义角色标注工具。
- 适合需要精确时空上下文的知识图谱构建与信息抽取任务。
我们提出一种针对场景上下文生成任务的神经架构:准确识别文本中提及事件或实体的相关时间与位置。上下文化信息抽取有助于限定自动化信息聚合在知识图谱中的有效性。该方法利用高质量标注的流行病学论文语料库中的时间与位置标注数据,训练编码器-解码器架构,并探索了训练过程中的数据增强技术。研究发现,经过微调的相对小型编码器-解码器模型,在预测特定实体或事件的相关情景信息方面,表现优于现成的大语言模型和语义角色标注解析器。
原文摘要 · Abstract (English)
We introduce a neural architecture finetuned for the task of scenario context generation: The relevant location and time of an event or entity mentioned in text. Contextualizing information extraction helps to scope the validity of automated finings when aggregating them as knowledge graphs. Our approach uses a high-quality curated dataset of time and location annotations in a corpus of epidemiology papers to train an encoder-decoder architecture. We also explored the use of data augmentation techniques during training. Our findings suggest that a relatively small fine-tuned encoder-decoder model performs better than out-of-the-box LLMs and semantic role labeling parsers to accurate predict the relevant scenario information of a particular entity or event.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。