arXiv:2608.25017cs.LG2026-08

用滚动解码重建提升潜空间模型的长时预测能力

Rollout-Decoded Reconstruction for Long-Horizon Prediction in Latent World Models

  • 训练时模拟推理过程,对每一步潜变量进行解码并惩罚误差
  • 在柯尔莫哥洛夫方程上预测有效时间提升1.8倍,达6.97单位
  • 无需额外参数,适合追求长时预测精度的研究者

潜空间世界模型在训练时以观测锚定的潜变量为解码目标,但在推理时则对自身自由运行的潜变量序列进行解码,存在偏差。滚动解码重建(RDR)通过单一损失项,在训练阶段就让模型以与评估完全相同的自由运行方式执行,对每个滚动潜变量进行解码,并基于真实值惩罚重建误差。该方法不增加参数,仅消耗训练计算开销,且当权重为零时退化为标准目标,所有对比均为单变量A/B测试。在混沌的柯尔莫哥洛夫-西瓦辛斯基方程上,RDR将有效预测时间(首次归一化误差达到0.5的时间)从3.87±0.23提升至6.97±0.42时间单位,参数量保持193,568不变,提升1.80倍;在未参与选择的种子和10/10预注册配置中均验证成功,增益比为1.71–2.50倍。结果来自单一系统,潜变量宽度增大的实验表明优势随其增长,两项经典任务的对照实验为初步尝试。

原文摘要 · Abstract (English)

A latent world model trains its decoder on latents anchored to observations, then deploys it on the model's own free-running rollout, hundreds of steps past the last observation. Rollout-Decoded Reconstruction (RDR) closes this gap with a single loss term that free-runs the model during training exactly as evaluation will, decodes every rollout latent, and penalizes reconstruction error against ground truth. The term adds no parameters, costs training-time compute only, and reduces to the standard objective at weight zero, so every comparison in this paper is a one-flag A/B. On the chaotic Kuramoto-Sivashinsky equation, RDR raises valid prediction time (the time to first crossing of normalized error 0.5) from $3.87 \pm 0.23$ to $6.97 \pm 0.42$ time units at an identical 193,568 parameters, a $1.80\times$ improvement confirmed on seeds never used in selection and holding in 10 of 10 preregistered configurations at ratios of 1.71-2.50$\times$. The results come from a single system; a sweep in which the advantage grows with latent width is descriptive, and control experiments on two classic tasks are preliminary.

潜空间模型长时预测生成建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。