arXiv:2512.12654cs.CLcs.LG2025-12

用角色互动图分析乌尔都小说叙事结构,识别作者风格。

Modeling Authorial Style in Urdu Novels Using Character Interaction Graphs and Graph Neural Networks

  • 将小说转为角色共现图,用图神经网络捕捉叙事结构
  • 在7位作者52部小说上达85.7%作者识别准确率
  • 为低资源语言文学分析提供新方法,适合文本风格研究者

传统作者分析多依赖词汇与风格特征,而对叙事结构的关注不足,尤其在乌尔都语等低资源语言中更为欠缺。本文提出一种基于图的框架,将乌尔都小说建模为角色互动网络,探索仅通过叙事结构能否推断作者风格。每部小说被表示为一个图,节点为角色,边代表角色在叙事邻近中的共现。系统比较了多种图表示方式,包括全局结构特征、节点级语义摘要、无监督图嵌入及有监督图神经网络。在包含52部乌尔都小说、7位作者的数据集上,学习得到的图表示显著优于手工设计和无监督基线,在严格的作者感知评估协议下最高达到0.857的准确率。

原文摘要 · Abstract (English)

Authorship analysis has traditionally focused on lexical and stylistic cues within text, while higher-level narrative structure remains underexplored, particularly for low-resource languages such as Urdu. This work proposes a graph-based framework that models Urdu novels as character interaction networks to examine whether authorial style can be inferred from narrative structure alone. Each novel is represented as a graph where nodes correspond to characters and edges denote their co-occurrence within narrative proximity. We systematically compare multiple graph representations, including global structural features, node-level semantic summaries, unsupervised graph embeddings, and supervised graph neural networks. Experiments on a dataset of 52 Urdu novels written by seven authors show that learned graph representations substantially outperform hand-crafted and unsupervised baselines, achieving up to 0.857 accuracy under a strict author-aware evaluation protocol.

叙事分析图神经网络低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。