arXiv:2604.20221cs.CL2026-04被引 1

用马尔可夫模型分析《叶甫盖尼·奥涅金》的语音结构,发现俄语原文与意大利语译文的节奏差异。

Markov reads Pushkin, again: A statistical journey into the poetic world of Evgenij Onegin

论文配图:Markov reads Pushkin, again: A statistical journey into the poetic world of Evgenij Onegin
图 1 · 摘自论文原文
  • 构建四状态马尔可夫链,捕捉音节的局部依赖与长程模式。
  • 俄语文本记忆深度逐渐下降,译文则保持稳定。
  • 通过语音探针揭示文字形式与主题发展的微妙关联。

本研究采用符号时间序列分析与马尔可夫建模,探索《叶甫盖尼·奥涅金》的语音结构,基于图示元音/辅音(V/C)编码,并对比一种当代意大利语译本。受马尔可夫原始方案启发,采用二进制编码构建极简概率模型,同时刻画局部V/C依赖与大尺度序列模式。一个四状态马尔可夫链被证明具有描述性准确性和生成能力,能复现原始序列的关键特征,如自相关性和记忆深度。所有发现均为探索性,旨在揭示结构性规律并提出关于叙事动力的假设。分析显示俄语文本与意语文本存在显著不对称:原文呈现记忆深度逐步降低,而译文维持较均匀分布。为进一步探究此差异,引入语音探针——短符号模式,连接表层结构与叙事线索。追踪文本展开过程中的探针变化,揭示了俄语文本中图示形式与主题发展之间的细微关联。通过重访马尔可夫对文学文本进行符号分析的初衷,并结合当代计算统计与数据科学工具,本研究表明,即使是最小化马尔可夫模型也可支持复杂诗学材料的探索性分析。当辅以粗粒度语言标注层时,此类模型为比较诗学提供通用框架,并证明风格化结构模式仍可通过植根于语言形式的简单表征保持可解析性。

原文摘要 · Abstract (English)

This study applies symbolic time series analysis and Markov modeling to explore the phonological structure of Evgenij Onegin-as captured through a graphemic vowel/consonant (V/C) encoding-and one contemporary Italian translation. Using a binary encoding inspired by Markov's original scheme, we construct minimalist probabilistic models that capture both local V/C dependencies and large-scale sequential patterns. A compact four-state Markov chain is shown to be descriptively accurate and generative, reproducing key features of the original sequences such as autocorrelation and memory depth. All findings are exploratory in nature and aim to highlight structural regularities while suggesting hypotheses about underlying narrative dynamics. The analysis reveals a marked asymmetry between the Russian and Italian texts: the original exhibits a gradual decline in memory depth, whereas the translation maintains a more uniform profile. To further investigate this divergence, we introduce phonological probes-short symbolic patterns that link surface structure to narrative-relevant cues. Tracked across the unfolding text, these probes reveal subtle connections between graphemic form and thematic development, particularly in the Russian original. By revisiting Markov's original proposal of applying symbolic analysis to a literary text and pairing it with contemporary tools from computational statistics and data science, this study shows that even minimalist Markov models can support exploratory analysis of complex poetic material. When complemented by a coarse layer of linguistic annotation, such models provide a general framework for comparative poetics and demonstrate that stylized structural patterns remain accessible through simple representations grounded in linguistic form.

诗歌分析马尔可夫模型语言结构比较诗学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。