arXiv:2608.17096cs.CL2026-08

破解《伏尼契手稿》关键:其字符、词与空格均非字母、单词或分隔符。

A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not

论文配图:A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not
图 1 · 摘自论文原文
  • 通过对照文本与伪文本验证,发现手稿单位并非字母、词或空格
  • 词间顺序信息极弱,仅占0.1%熵值,远低于对照组的2-10%
  • 词边缘字符有显著关联(0.2比特互信息),暗示非传统语言结构

《伏尼契手稿》(Beinecke MS 408)通常基于三个未明言假设分析:其符号为字母,空白之间的字符串为词,每个空白为词间分隔。本文以Zandbergen-Landini转写为基础,结合真实文本、密码文本与伪文本对照,并进行书页级重采样检验,发现三者均不成立。失败模式一致:手稿中字符序列主要分布在词边缘及词间渐变边界,而非词内连续。符号规律性过强(条件熵2.7比特),远高于拉丁语、意大利语和英语(约3.5比特),表明其反映的是稳定存在的多符号单元结构,而非单对一替换。尽管词构成合理词汇,但一个词预测下一个词的信息量不足1%熵值,低于所有对照组(2-10%)。词边字符间互信息达0.2比特,高于所有真实文本对照。空白分为两类:标记不确定的分隔符行为如词内连接点,物理宽度更窄(图像坐标区分度AUC 0.905),且在去除所有空格后仍被学习单元跨越。该特征可有效区分真手稿。已发表的仿手稿密码与自引生成器虽复现低熵、单元尺度与弱词序,却无法再现边缘字符耦合与开放、罕见词占比高的词汇(70%为独有类型,而对照组为41%与59-60%)。因此,任何解释都必须建立在实证基础上,而非预设字符、词、分隔符对应字母、词、词空间,这些测量结果才是评估标准。

原文摘要 · Abstract (English)

The Voynich manuscript (Beinecke MS 408) is usually analysed on three unstated assumptions: that its glyphs are letters, that the strings between blanks are words, and that every blank is a word space. We test all three against the Zandbergen-Landini transliteration with matched prose, cipher, and pseudo-text controls and quire-level resampling. None holds, and the failures share a shape: the order in Voynichese sits at the edges of tokens and at graded boundaries between them, not in the succession of tokens themselves. Glyph regularity is too strong for one-to-one substitution of any tested plaintext (conditional entropy 2.7 bits against about 3.5 for Latin, Italian, and English) and resolves instead onto a quire-stable scale of recurrent multi-symbol units. Tokens form a plausible vocabulary, yet the identity of one token predicts the next by under 1% of token entropy, below every matched control (2-10%), while the glyphs at token edges share 0.2 bits of mutual information, more than in any prose control. Blanks fall into two regimes: the separators transcribers marked uncertain behave like word-internal junctures, are physically narrower on the page (AUC 0.905 from independent image coordinates, with the same sign in a small blind ink audit), and are crossed by learned units even when every space is erased before learning. This profile is also what discriminates. A published Voynich-imitating cipher and a self-citation text generator both reproduce the low entropy, the unit scale, the weak token order, and the null result of a calibrated substitution attack; neither reproduces the edge-glyph coupling or the open, hapax-rich vocabulary (70% singleton types against 41% and 59-60%). Any account of the manuscript must therefore earn, rather than assume, the step from glyphs, tokens, and separators to letters, words, and word spaces, and these are the measurements on which to do so.

手稿破解语言模型信息熵符号分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。