arXiv:2503.15681cs.IR2025-03被引 11

用深度模型语义构建故事链,自动提取连贯叙事。

Narrative Trails: A Method for Coherent Storyline Extraction via Maximum Capacity Path Optimization

  • 基于深度模型嵌入构建稀疏连贯图,优化最大容量路径。
  • 在两个任务上验证,可扩展至大规模文本数据集。
  • 无需复杂规则,适合通用叙事提取场景。

传统信息检索侧重从大数据集中寻找相关文档,但未对结果进行结构化组织。而以叙事形式组织信息——即有序文档序列构成连贯故事线——有助于揭示数据中观点间的关联与关系。尽管叙事结构意义重大,现有算法提取方法仍稀缺,多数依赖复杂的词级启发式规则和辅助文档结构,且难以扩展至大规模或通用场景。本文提出Narrative Trails,一种高效、通用的大型文本语料叙事提取方法。该方法利用深度学习模型的潜在空间语义信息,构建稀疏连贯图,并通过最大化故事线最小连贯性来提取叙事。在两个不同叙事提取任务上的定量评估表明,Narrative Trails具备良好的泛化性和可扩展性,同时简化了提取流程。

原文摘要 · Abstract (English)

Traditional information retrieval is primarily concerned with finding relevant information from large datasets without imposing a structure within the retrieved pieces of data. However, structuring information in the form of narratives--ordered sets of documents that form coherent storylines--allows us to identify, interpret, and share insights about the connections and relationships between the ideas presented in the data. Despite their significance, current approaches for algorithmically extracting storylines from data are scarce, with existing methods primarily relying on intricate word-based heuristics and auxiliary document structures. Moreover, many of these methods are difficult to scale to large datasets and general contexts, as they are designed to extract storylines for narrow tasks. In this paper, we propose Narrative Trails, an efficient, general-purpose method for extracting coherent storylines in large text corpora. Specifically, our method uses the semantic-level information embedded in the latent space of deep learning models to build a sparse coherence graph and extract narratives that maximize the minimum coherence of the storylines. By quantitatively evaluating our proposed methods on two distinct narrative extraction tasks, we show the generalizability and scalability of Narrative Trails in multiple contexts while also simplifying the extraction pipeline.

叙事提取连贯性深度模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。