arXiv:2605.28643cs.CL2026-05被引 1

将人物互动与文本语境结合,用动态图学习文学表征。

GraphLit: Learning Text-Enriched Dynamic Character Network Representations for Literary Study

论文配图:GraphLit: Learning Text-Enriched Dynamic Character Network Representations for Literary Study
图 1 · 摘自论文原文
  • 构建时序局部异质图,将人物与其上下文绑定
  • 在12项任务中超越纯文本与纯图基线,提升显著
  • 适合研究叙事非线性与人物关系动态的学者

现有方法将文学文本表示为图或图序列时,多聚焦人物互动,忽略其所在文本语境。本文提出动态异质人物网络(DHCNs),将长篇小说拆分为时间局部化的异质图,使人物与其文本上下文对齐。从古腾堡计划中提取约20,000个DHCNs,设计自监督学习框架GraphLit,通过掩码图自编码目标学习丰富的文学表征。在12项人物相关任务中,GraphLit显著优于仅用文本、仅用图及先前混合基线。消融实验表明,将人物锚定于语境是性能提升主因,而显式编码叙事顺序与人物关系则带来任务依赖性改善。最后,展示了DHCNs与GraphLit在文学分析中的应用,揭示叙事非线性与动态社会特征之间的关联。

原文摘要 · Abstract (English)

Methods to represent literary texts as graphs or sequences of graphs mainly focus on representing character interactions, and often overlook another crucial aspect: the textual context in which characters interact. We introduce Dynamic Heterogeneous Character Networks (DHCNs), which organize long novels into temporally localized heterogeneous graphs that align characters with their textual contexts. We extract around 20,000 DHCNs from Project Gutenberg, and propose GraphLit, a self-supervised learning framework that learns rich literary representations through a masked graph autoencoder objective. Across a wide range of 12 character-related tasks, GraphLit improves over text-only, graph-only and prior hybrid baselines. Ablations over different kinds of dynamic graph structures and architectural elements show that grounding characters in their context is the main performance driver, while explicitly encoding narrative order and character relationships provide task-dependent improvements. Finally, we demonstrate the applicability of DHCNs and GraphLit for literary analysis by studying the link between narrative non-linearity and dynamic social features.

文学分析动态图自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。