用混合模型分析普希金诗体小说,找出核心主题及其叙事演变。
Hybrid topic modelling for computational close reading: Mapping narrative themes in Pushkin's Evgenij Onegin
- 结合LDA与sPLS-DA,从词频中挖掘主题并提升可解释性。
- 在35个文段中识别出5个稳定可读的主题,关键节点匹配叙事高潮。
- 适合文学分析者与计算人文研究者,方法透明且可复现。
本研究提出一种混合主题建模框架,用于计算文学分析,融合潜在狄利克雷分配(LDA)与稀疏偏最小二乘判别分析(sPLS-DA),以刻画叙事诗歌的主题结构与时间动态。以普希金的诗体小说《叶甫盖尼·奥涅金》的意大利语译本为案例,测试小语料库下无监督与有监督词汇结构是否收敛。文本被分割为35个文档,包含词形还原后的实词,最终得出5个稳定且可解释的主题。为应对小语料库不稳定性,采用多种子共识协议。使用sPLS-DA作为有监督探针,识别出能细化每个主题的词汇标记。叙事枢纽(即连续诗节构成的关键情节段落)将词袋模型拓展至叙事层面,揭示主题混合如何对应诗歌的情感与结构弧线。该框架并非取代传统文学解读,而是提供一种计算化的细读方式,证明轻量级概率模型可在忽略格律、音韵等特征的前提下,对复杂诗性叙事生成可复现的主题图谱。尽管依赖单一词形还原译本,该方法仍为其他高密度文学文本的比较研究提供了透明的方法模板。
原文摘要 · Abstract (English)
This study presents a hybrid topic modelling framework for computational literary analysis that integrates Latent Dirichlet Allocation (LDA) with sparse Partial Least Squares Discriminant Analysis (sPLS-DA) to model thematic structure and longitudinal dynamics in narrative poetry. As a case study, we analyse Evgenij Onegin-Aleksandr S. Pushkin's novel in verse-using an Italian translation, testing whether unsupervised and supervised lexical structures converge in a small-corpus setting. The poetic text is segmented into thirty-five documents of lemmatised content words, from which five stable and interpretable topics emerge. To address small-corpus instability, a multi-seed consensus protocol is adopted. Using sPLS-DA as a supervised probe enhances interpretability by identifying lexical markers that refine each theme. Narrative hubs-groups of contiguous stanzas marking key episodes-extend the bag-of-words approach to the narrative level, revealing how thematic mixtures align with the poem's emotional and structural arc. Rather than replacing traditional literary interpretation, the proposed framework offers a computational form of close reading, illustrating how lightweight probabilistic models can yield reproducible thematic maps of complex poetic narratives, even when stylistic features such as metre, phonology, or native morphology are abstracted away. Despite relying on a single lemmatised translation, the approach provides a transparent methodological template applicable to other high-density literary texts in comparative studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。