区分未来分支是因观测模糊还是动态随机,提升模型可解释性。
Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models
- 通过重复扰动识别状态与噪声的交互关系,分离观测模糊与真实随机。
- 在多个场景下实现90%以上路由准确率,优于仅依赖预测质量的方法。
- 适合需要理解模型决策依据的研究者,如安全关键系统设计。
校准的随机世界模型能揭示未来不确定性程度,但无法说明其来源。相同的条件未来规律可能源于观测对物理状态的混淆,或在确定完整状态后动力学仍保持随机。我们证明,即使使用完美的概率预测器,普通转移过程也无法区分这两种原因。ClosurePairs通过交叉兼容微观状态与重复外源扰动,估计状态、噪声及其交互方差,使二者可被识别。核心结论是:在有限分层采样下,预测难度决定有效计算规模,而别态/过程构成提供互补信息——指示应解析当前状态还是采样未来随机性。ClosurePairs在不改变似然的前提下恢复源归属,在非线性交互基准中降低等预算分解误差,并支持仅观察路由。在精确边缘分布的MetaWorld双胞胎上,仅输出分配器处于随机水平,而基于关闭监督的探针在冻结JEPA-WM特征上实现89.8%-100%路由准确率;在独立ManiSkill PushCube验证中,随机RSSM的输出与隐状态均处于随机水平,而仅使用RGB的闭包探针在五次种子测试中实现100%路由,无论是否为内部或几何/相机领域外样本,匹配直接分配而非超越。在五个未见分配菜单中,同一闭包探针在内部/外部环境下分别实现92.5%/90.4%准确率,无需新标签,而冻结直接分配器仅为37.9%/32.9%。因此,ClosurePairs是一种不可从预测质量中恢复的可识别、可重用机制目标。
原文摘要 · Abstract (English)
A calibrated stochastic world model can reveal how uncertain a future is without revealing why it branches. The same conditional future law can arise because an observation aliases physical states or because dynamics remain random after the declared full state is fixed. We prove that ordinary transitions cannot identify these two sources, even for a perfect probabilistic predictor. ClosurePairs makes them identifiable by crossing compatible microstates with repeated exogenous disturbances and estimating state, noise, and state-noise interaction variance. The central consequence is operational: under finite hierarchical sampling, forecast difficulty governs the useful compute scale, while the alias/process composition provides complementary information about its direction-resolving the current state or sampling future randomness. ClosurePairs recovers source attribution at unchanged likelihood, reduces equal-budget decomposition error in a nonlinear interaction benchmark, and supports observation-only routing. On exact-marginal MetaWorld twins, an output-only allocator is at chance while a Closure-supervised probe on frozen JEPA-WM features routes 89.8-100%. In an independent ManiSkill PushCube confirmation, a stochastic RSSM's outputs and latents remain at chance, whereas an RGB-only Closure probe routes 100% under both ID and geometry/camera OOD over five seeds, matching direct allocation rather than exceeding it. Across five unseen allocation menus, the same Closure probe routes 92.5%/90.4% ID/OOD with no new oracle labels, versus 37.9%/32.9% for a frozen direct allocator. ClosurePairs is therefore an identifiable, reusable mechanism target that cannot be recovered from forecast quality alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。