用概率机器学习分析柯拉茨序列停止时间,揭示其分布规律与模结构影响。
Bayesian Modeling of Collatz Stopping Times: A Probabilistic Machine Learning Perspective
- 构建贝叶斯分层负二项回归模型,基于对数规模和模8余数预测停止时间。
- 在10^7以内数据上,模型预测似然显著优于生成式方法,误差率更低。
- 发现模8结构是导致停止时间异质性的关键因素,适合数论与概率建模研究者。
我们从概率机器学习视角研究了不超过10^7的柯拉茨总停止时间τ(n)。实证显示,τ(n)为偏斜且高度过度分散的计数变量,具有明显的算术异质性。我们提出了两种互补模型:第一,基于简单协变量(log n 和 n mod 8)的贝叶斯分层负二项回归(NB2-GLM),通过后验及后验预测分布量化不确定性;第二,基于奇数块分解的机制生成近似模型:对奇数m,将3m+1表示为2^{K(m)}m',其中m'为奇数,K(m)=v_2(3m+1)≥1;通过随机化块长并以狄利克雷-多项式更新校准。在留出数据上,NB2-GLM的预测似然显著高于奇数块生成器。将块长分布条件于m mod 8后,生成器的分布拟合明显改善,表明低阶模结构是τ(n)异质性的主要驱动因素。
原文摘要 · Abstract (English)
We study the Collatz total stopping time $τ(n)$ over $n\le 10^7$ from a probabilistic machine learning viewpoint. Empirically, $τ(n)$ is a skewed and heavily overdispersed count with pronounced arithmetic heterogeneity. We develop two complementary models. First, a Bayesian hierarchical Negative Binomial regression (NB2-GLM) predicts $τ(n)$ from simple covariates ($\log n$ and residue class $n \bmod 8$), quantifying uncertainty via posterior and posterior predictive distributions. Second, we propose a mechanistic generative approximation based on the odd-block decomposition: for odd $m$, write $3m+1=2^{K(m)}m'$ with $m'$ odd and $K(m)=v_2(3m+1)\ge 1$; randomizing these block lengths yields a stochastic approximation calibrated via a Dirichlet-multinomial update. On held-out data, the NB2-GLM achieves substantially higher predictive likelihood than the odd-block generators. Conditioning the block-length distribution on $m\bmod 8$ markedly improves the generator's distributional fit, indicating that low-order modular structure is a key driver of heterogeneity in $τ(n)$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。