arXiv:2504.05954cs.CLcs.LG2025-04被引 2

无需标注数据,自动将叙事文本中的地点轨迹映射到空间地图上。

Unsupervised Location Mapping for Narrative Corpora

  • 基于长上下文大模型,无监督构建叙事文本中的地点地图。
  • 在大屠杀证词与湖区文学中验证,定位准确率显著优于基线。
  • 适合对历史叙事、文学地理分析感兴趣的学者与研究者。

本文提出无监督地点映射任务,旨在将单篇叙事的事件轨迹映射到由大量叙事文本共同涉及的地点空间地图上。尽管该任务具有基础性和普适性,但相关研究极少。任务包含两部分:(1) 从多篇文本中推导出地点构成的‘地图’;(2) 从单个叙事中提取事件轨迹并定位到地图上。利用近期大语言模型上下文长度的提升,我们提出完全无监督的流水线方法,无需预先定义标签集。我们在两个不同领域进行测试:(1) 大屠杀证词;(2) 英格兰湖区文学(跨世纪旅行文学)。通过内在与外在评估,结果令人鼓舞,为该任务建立了基准与评估范式,并揭示了当前挑战。

原文摘要 · Abstract (English)

This work presents the task of unsupervised location mapping, which seeks to map the trajectory of an individual narrative on a spatial map of locations in which a large set of narratives take place. Despite the fundamentality and generality of the task, very little work addressed the spatial mapping of narrative texts. The task consists of two parts: (1) inducing a ``map'' with the locations mentioned in a set of texts, and (2) extracting a trajectory from a single narrative and positioning it on the map. Following recent advances in increasing the context length of large language models, we propose a pipeline for this task in a completely unsupervised manner without predefining the set of labels. We test our method on two different domains: (1) Holocaust testimonies and (2) Lake District writing, namely multi-century literature on travels in the English Lake District. We perform both intrinsic and extrinsic evaluations for the task, with encouraging results, thereby setting a benchmark and evaluation practices for the task, as well as highlighting challenges.

无监督学习叙事分析地理信息大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。