arXiv:2604.09234cs.LGcs.AI2026-04被引 1

古人八卦序列看似有学习优势,实则让神经网络训练更差。

Statistical Properties of the King Wen Sequence: An Anti-Habituation Structure That Does Not Improve Neural Network Training

  • 用蒙特卡洛分析64个卦象排列的统计特性
  • 该序列在四方面显著优于随机:跨距大、自相关负、阳爻均衡、组内组间不对称
  • 实验表明其会破坏训练稳定性,不适合做课程学习

《易经》中周文王卦序将64个六维二进制状态按特定顺序排列,三千年未解。本文通过10万次蒙特卡洛置换分析,发现该序列具有四项统计显著特征:转换距离高于98.2百分位、一阶自相关为负(p=0.037)、每组四个卦象阳爻平衡(p=0.002)、组内与组间距离不对称(99.2百分位)。这些特征看似符合课程学习与好奇心驱动探索原则,故假设可能提升神经网络训练。我们通过三个实验验证:学习率调制、课程排序、种子敏感性分析,在NVIDIA RTX 2060(PyTorch)与Apple Silicon(MLX)平台进行。结果一致负面:文王序列学习率调制在所有振幅下性能下降;作为课程排序时,在一平台为最差非连续排序,另一平台与噪声相当;30次种子测试显示,仅文王序列的退化超过自然种子方差。原因在于其高方差——正是使其统计独特的根源——会破坏梯度优化稳定性。固定组合序列的反习惯化不等于有效训练动态。

原文摘要 · Abstract (English)

The King Wen sequence of the I-Ching (c. 1000 BC) orders 64 hexagrams -- states of a six-dimensional binary space -- in a pattern that has puzzled scholars for three millennia. We present a rigorous statistical characterization of this ordering using Monte Carlo permutation analysis against 100,000 random baselines. We find that the sequence has four statistically significant properties: higher-than-random transition distance (98.2nd percentile), negative lag-1 autocorrelation (p=0.037), yang-balanced groups of four (p=0.002), and asymmetric within-pair vs. between-pair distances (99.2nd percentile). These properties superficially resemble principles from curriculum learning and curiosity-driven exploration, motivating the hypothesis that they might benefit neural network training. We test this hypothesis through three experiments: learning rate schedule modulation, curriculum ordering, and seed sensitivity analysis, conducted across two hardware platforms (NVIDIA RTX 2060 with PyTorch and Apple Silicon with MLX). The results are uniformly negative. King Wen LR modulation degrades performance at all tested amplitudes. As curriculum ordering, King Wen is the worst non-sequential ordering on one platform and within noise on the other. A 30-seed sweep confirms that only King Wen's degradation exceeds natural seed variance. We explain why: the sequence's high variance -- the very property that makes it statistically distinctive -- destabilizes gradient-based optimization. Anti-habituation in a fixed combinatorial sequence is not the same as effective training dynamics.

机器学习统计分析神经网络易经

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。