比较中英文文学网络的平均最短路径,发现标点影响中文网络连通性。
Average shortest-path length in word-adjacency networks: Chinese versus English
- 将标点视为词语构建语义网络,分析中英文文学作品结构。
- 含标点时中英文网络的路径长度随规模增长趋势相似。
- 忽略标点会使中文网络路径显著变长,凸显标点作用。
复杂网络为分析自然语言等系统的复杂结构提供了有力工具。本文研究了从不同时期中英文文学作品构建的动态词邻接网络的拓扑结构。不同于传统仅考虑词汇,本文还将标点符号视为普通词语。这一做法基于两点:(1) 标点传递情感状态信息,辅助逻辑分组与阅读停顿,减少歧义;(2) 前期研究表明,标点在齐夫分析中表现如词汇,与词汇结合可提升作者识别效果。我们聚焦不同年代及单部小说原文与译文的平均最短路径长度 $L(N)$ 随网络规模 $N$ 的函数关系。通过拟合生长网络模型,实证结果与模型吻合良好。结果显示:当包含标点时,中英文 $L(N)$ 渐近行为相似;若忽略标点,中文 $L(N)$ 显著增大。
原文摘要 · Abstract (English)
Complex networks provide powerful tools for analyzing and understanding the intricate structures present in various systems, including natural language. Here, we analyze topology of growing word-adjacency networks constructed from Chinese and English literary works written in different periods. Unconventionally, instead of considering dictionary words only, we also include punctuation marks as if they were ordinary words. Our approach is based on two arguments: (1) punctuation carries genuine information related to emotional state, allows for logical grouping of content, provides a pause in reading, and facilitates understanding by avoiding ambiguity, and (2) our previous works have shown that punctuation marks behave like words in a Zipfian analysis and, if considered together with regular words, can improve authorship attribution in stylometric studies. We focus on a functional dependence of the average shortest path length $L(N)$ on a network size $N$ for different epochs and individual novels in their original language as well as for translations of selected novels into the other language. We approximate the empirical results with a growing network model and obtain satisfactory agreement between the two. We also observe that $L(N)$ behaves asymptotically similar for both languages if punctuation marks are included but becomes sizably larger for Chinese if punctuation marks are neglected.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。