arXiv:2412.07919q-bio.NCcs.CL2024-12被引 2

发现意大利语文本中词语分布符合玻色-爱因斯坦统计

Identifying Quantum Mechanical Statistics in Italian Corpora

  • 用自研理论框架分析意大利语文本词频分布
  • 所有文本均显示词分布符合玻色-爱因斯坦统计,偏离麦克斯韦-玻尔兹曼统计
  • 词随机化模拟温度升高,破坏语义相干性,使经典统计占优

我们对人类语言生成文本中的词汇统计行为进行了理论与实证研究。通过分析来自特定文学语料库的多种意大利语文本,首先推广了自研理论框架以识别大文本中的‘量子统计’特征。结果表明,在所有分析文本中,词汇分布均符合‘玻色-爱因斯坦统计’,显著偏离‘麦克斯韦-玻尔兹曼统计’。随后引入‘词随机化’效应,发现该过程使两种统计模型差异减弱。这些结果验证了此前在英语文本中观察到的模式,强烈暗示相同词汇因语义关联而‘聚集’,可解释为由‘上下文更新’引发的‘量子纠缠’效应。词随机化可类比为温度升高,破坏语义相干性,使经典统计取代量子统计。最后,文章探讨了物理学中量子统计的起源。

原文摘要 · Abstract (English)

We present a theoretical and empirical investigation of the statistical behaviour of the words in a text produced by human language. To this aim, we analyse the word distribution of various texts of Italian language selected from a specific literary corpus. We firstly generalise a theoretical framework elaborated by ourselves to identify 'quantum mechanical statistics' in large-size texts. Then, we show that, in all analysed texts, words distribute according to 'Bose--Einstein statistics' and show significant deviations from 'Maxwell--Boltzmann statistics'. Next, we introduce an effect of 'word randomization' which instead indicates that the difference between the two statistical models is not as pronounced as in the original cases. These results confirm the empirical patterns obtained in texts of English language and strongly indicate that identical words tend to 'clump together' as a consequence of their meaning, which can be explained as an effect of 'quantum entanglement' produced through a phenomenon of 'contextual updating'. More, word randomization can be seen as the linguistic-conceptual equivalent of an increase of temperature which destroys 'coherence' and makes classical statistics prevail over quantum statistics. Some insights into the origin of quantum statistics in physics are finally provided.

语言统计量子力学语义凝聚

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。