为具有对称性的世界模型提供可信赖的预测时长认证方法。
Conformal Orbit-Valid Trust Horizons for Equivariant World Models

- 基于群对称性与残差估计,构建校准后的信任时长曲线。
- 50次审计中零违规,违反率上限95%置信度下为5.8%。
- 适用于需要安全预测时长验证的机器人与强化学习场景。
学习型世界模型仅在滚动误差可控的时间范围内有效。本文研究具有已知群对称性的潜在世界模型的信任时长认证问题。给定一步潜在残差和有限时间扩展估计,构造原始时长曲线,并通过分裂-容错乘性因子进行校准。在可重复审计集上,容错因子γ_α=1.0:原始证书已足够保守。50次稳定审计中未出现反保守违规,对应精确二项式95%置信上界为5.8%的违规率。主要结构结果表明,精确等变性可将校准后的信任时长曲线沿群轨道传输:当环境动力学、编码器、预测器、动作变换及潜在度量满足指定等变/不变条件时,滚动误差与信任时长在轨道上保持恒定。实证结果显示,所实现模型的轨道传输残差较小,14次轨道审计中中位数为1.1%,最大值为4.1%。证书非退化(中位认证时长与实测时长比为0.67)。证书级校准成本分析显示两种互补范式:在对称2D基底上,等变、普通与增强模型均仅需单个校准区域即达轨道有效——无区分,因基底本身使非等变基线近似轨道鲁棒;而在3D偏航审计中,等变模型获得单区域安全且非退化的轨道有效证书,而健康非等变基线则面临违规、松弛、锐度或额外区域成本。该证书为保守分布审计而非全局可达性保证,当前3D CEM-MPC行为层中未证实证书引导子目标间距的有效性。
原文摘要 · Abstract (English)
Learned world models are useful only over horizons on which their rollout error remains controlled. We study trust-horizon certification for latent world models with known group symmetries. Given a one-step latent residual and a finite-time expansion estimate, we form a raw horizon curve and calibrate it with a split-conformal multiplicative factor. On the reproducible audit set, the conformal factor is $γ_α=1.0$: the raw certificate is already conservative under the audit protocol. Across 50 stable audits, we observe zero anti-conservative violations, corresponding to an exact-binomial 95% upper bound of 5.8% on the violation rate. Our main structural result is that exact equivariance transports a calibrated trust-horizon curve over the group orbit: when the environment dynamics, encoder, predictor, action transform, and latent metric satisfy the stated equivariance/invariance conditions, rollout errors and trust horizons are orbit-constant. Empirically, the implemented models exhibit small orbit-transport residuals, with median 1.1% and maximum 4.1% over 14 orbit audits. The certificate is also non-vacuous (median certified-to-measured horizon ratio 0.67). A certificate-level calibration-cost study shows two complementary regimes. On a symmetric 2D substrate, equivariant, plain, and augmented models are all orbit-valid from a single calibration sector -- no separation, because the substrate already makes non-equivariant baselines approximately orbit-robust. A 3D yaw audit shows the other regime: the equivariant model obtains a one-sector safe and non-vacuous orbit-valid certificate, while healthy non-equivariant baselines pay violation, slack, sharpness, or additional-sector cost. The certificate is a conservative, distributional audit rather than a global reachability guarantee, and certificate-guided subgoal spacing is not confirmed in the current 3D CEM-MPC behavior layer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。