中文字符的简繁改革提升了编码效率,使常用字更短。
Zipf's Law of Abbreviation in a Logographic Script: Coding-Theoretic Bounds on Chinese Character Stroke Counts
- 以笔画为代价单位,量化汉字编码压缩程度
- 简化改革使常用字平均笔画从12.71降至7.22,优化度提升至0.668
- 笔画结构非唯一可解,二维布局增强可读性,牺牲部分压缩
Zipf简写规律——高频词趋向于简短——是语言中最具支持性的规律之一。近期研究从验证该规律转向测量词汇库相对于理论基准的压缩程度。此前研究集中于拼音和音节文字,本文将其拓展至表意文字,以笔画为发音成本单位,汉字为编码形式。结合涵盖20,902个CJK基本字符的笔顺数据库及两个独立语料库(分别含2.589亿和1.933亿词符),发现字符类型平均需12.71笔,而实际文本中平均仅7.22笔。采用Petrini等人(2026)的双重归一化最优性评分,简化字库存达到Ω = 0.668,复现语料库为0.609,均位于其报告的20种语言8种书写系统的62%-67%区间内,表明压缩上限与书写系统类型和成本单位无关。表意文字使绝对编码界限可计算:五元分类笔画下,5进制霍夫曼最优为4.34笔,熵界为4.28笔,观测系统达1.66倍最优值。该差距并非冗余,而是结构使然:5^(-l_i)的Kraft和在频率列表为2.05,全库为5.03,表明笔画序列无法在一维上唯一解码;汉字通过二维笔画布局实现歧义消除,而非序列顺序,由此换取了部件透明性。将20世纪中期的简化改革视为可控压缩事件,发现其将最优性从0.555提升至0.668,节省主要集中在最常用的1000个字符。
原文摘要 · Abstract (English)
Zipf's law of abbreviation -- the tendency of frequent forms to be short -- is one of the best-supported regularities in language, and recent work has moved from demonstrating it to measuring how far lexicons are compressed relative to principled baselines. That programme has so far addressed word lengths in alphabetic and syllabic scripts. We transfer it to a logographic script, taking the stroke as the unit of articulatory cost and the Chinese character as the coded form. Combining a stroke-order database covering all 20,902 characters of the CJK basic block with two independent frequency corpora (258.9M and 193.3M tokens), we find that the mean character type costs 12.71 strokes but the mean character token in running text only 7.22. Using the dually normalised optimality score of Petrini et al. (2026), the simplified inventory reaches Omega = 0.668, with the replication corpus at 0.609 -- inside and just below the 62-67% band those authors report for word lengths across 20 languages and 8 scripts, suggesting a compression ceiling largely independent of script type and cost unit. A logographic script also makes absolute coding bounds computable, since strokes come from a closed five-element taxonomy: the exact 5-ary Huffman optimum is 4.34 strokes and the entropy bound 4.28, so the observed system is 1.66x above optimal coding. This gap is not slack but structure. The Kraft sum of 5^(-l_i) is 2.05 on the frequency list and 5.03 on the full inventory, so stroke strings are provably not uniquely decodable in one dimension; characters are disambiguated by the two-dimensional arrangement of strokes, not their sequence, and the forgone compression buys componential transparency. Finally, treating the mid-twentieth-century simplification reform as a controlled compression event, we find it raised optimality from 0.555 to 0.668, with savings concentrated in the 1,000 commonest characters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。