用合成数据对比,发现印度河文字既不像语言也不像纹章或编码系统。
How Non-Linguistic Is the Indus Sign System? A Synthetic-Baseline Scorecard

- 构建两种非语言基线:纹章与行政编码,模拟频率和位置规律。
- 1916条铭文在4项指标上介于两类基线之间,无一完全匹配。
- 结论支持其可能为语言,适合考古学与符号学研究者阅读。
印度河文字(约公元前2600-1900年)是否表意语言长期存在争议。本文提出多维度判别框架,将实证的印度河铭文与两类计算机生成的非语言基线对比:一类模仿纹章体系,另一类模仿行政编码体系,两者均基于六组已知非语言文本的齐普夫分布、位置约束与二元组依赖关系校准。评估四个核心属性:文本简短性、重复公式化短语、罕见词率及位置刚性,对应于Farmer-Sproat-Witzel(2004)的批评。对1,916条去重铭文(584个独特符号,11,110个字符)进行分析,结果表明印度河文本未完全契合任一基线,在四项指标中处于中间位置,无法被单一基线复现。同时与七组真实非语言文本(包括Sproat, 2014年数据集)比较,亦无一能完整复制其统计特征。关键结果复现了此前研究中的齐普夫斜率-1.49与条件熵3.23比特。所有代码与数据公开可得。
原文摘要 · Abstract (English)
Whether the Indus Valley sign system (c. 2600-1900 BCE) encodes spoken language has been debated for decades. This paper introduces a multi-metric discrimination framework that tests the observed Indus corpus against two kinds of computer-generated non-linguistic baseline -- one mimicking a heraldic emblem system, the other an administrative coding system -- each calibrated with Zipfian frequency distributions, positional constraints, and bigram dependencies derived from six attested non-linguistic corpora. The scorecard evaluates four properties central to the Farmer-Sproat-Witzel (2004) critique: text brevity, repeated formulaic phrases, hapax legomenon rate, and positional rigidity. Applying this framework to 1,916 deduplicated inscriptions (584 unique signs, 11,110 tokens) from the ICIT/Yajnadevam digitization, we find that the Indus corpus does not match either baseline cleanly. Across the four metrics examined, the Indus corpus occupies an intermediate position relative to the two baseline families, matching neither cleanly. Neither a heraldic nor an administrative generator can reproduce all four properties at once. We also compare against seven real-world non-linguistic corpora including Sproat's (2014) datasets, finding that no attested non-linguistic system reproduces the full Indus statistical profile either. We replicate key prior results including a Zipf slope of -1.49 and conditional entropy of 3.23 bits. All code and data are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。