arXiv:2608.28541cs.LGcs.AI2026-08

模型在不可达区域的错误无法被验证,但其危害取决于拓扑结构与可及性的关系。

An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models

  • 以通道宽度gamma为调节键,探索模型误差在三种状态间的转变
  • 通道可及性决定危险程度:可触通道使错误成本从1.09降至近0
  • 修复失败源于参数与传感器限制,需匹配错误维度才能有效缓解

一个通过采样门验证的代码世界模型可在门内完全正确,而在门外任意错误。我们分析了认证模型所能知晓的内容及其错误代价,当缺失部分为环形冻结模式时。门商数使问题精确化:确定性接受仅能确定可达查询集上的模型;超出可达范围即为规范选择。在最小环仪器上,我们证明了极端情况(一个无法被任何采样门验证且比特无害的填充圆盘伪影),并利用大语言模型合成,在三个模型族中测量单一控制参数(通道宽度gamma)如何驱动该伪影穿越三个阶段:不可验证且无害、可验证且有代价、立即被证伪。三条原则组织实证结果:第一,危险是相对于可达性的拓扑问题;规划者可用通道使盲模型的利用成本从1.09降至接近0(拐点gamma≈0.1),而具有相同第一贝蒂数的隐藏通道则维持高风险(1.12)。第二,修复受参数与传感器边界约束:任一族均无法从外部证据恢复该区域;内部模型虽能提出正确拓扑,但无法确定参数,所提拓扑跟踪的是引导持久同调摘要中的错误beta_1(传感器几何分辨率极限),而非真实值。第三,缓解措施必须匹配错误的维度和方向:点屏障对一维边界无效,维度匹配的持久屏障将利用降为两课时短暂事件(0.999至0.058),双自由度证书则对称地消除虚构模式失败(1.769至0.029)。在n维情况下,壳层使误识近乎必然,而危险仍完全可利用:两个轴独立。

原文摘要 · Abstract (English)

A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wrong beyond it. We characterize what a certified model can know, and what its errors can cost, when the omission is an annular freeze mode enclosing an unreachable interior. The gate quotient makes the question precise: acceptance-with-certainty determines the model exactly on the reachable query set; beyond reach is gauge. On a minimal ring instrument we prove the extreme case (a wrong-topology filled-disc artifact unfalsifiable by any sampling gate and bitwise harmless at play) and measure, with LLM synthesis across three model families, how one knob (a channel of width gamma) walks the same artifact through three regimes: unfalsifiable-and-harmless, falsifiable-and-costly, and instantly falsified. Three principles organize the empirics. First, danger is topology relative to reach: a channel the planner can use collapses the blind model's exploitation (play cost 1.09 to ~0 over a knee at gamma ~ 0.1), while a hidden channel with the same first Betti number keeps it at full strength (1.12). Second, repair is parameter-bound and sensor-bound: no family recovers the region from outside evidence; from inside, models pose the right topology but cannot pin its parameters, and the posed topology tracks the guiding persistent-homology summary's wrong beta_1 (a sensor with a measured geometric resolution limit), not the truth. Third, mitigation must match the error's dimension and direction: point fences fail against the one-dimensional boundary, a dimension-matched persisted fence collapses exploitation to a two-lesson transient (0.999 to 0.058), and the dual freedom certificate collapses the invented-mode failure symmetrically (1.769 to 0.029). In n dimensions the shell makes misidentification near-certain while the danger stays fully exploitable: the two axes are independent.

拓扑学习模型可信误差分析持续同调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。