arXiv:2608.14650cs.LGcs.RO2026-08

用精确重置实验验证预测接口能否有效降低决策成本

Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade

论文配图:Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade
图 1 · 摘自论文原文
  • 构建配对评估框架,统一初始状态与动作集测试中/全模型切换效果
  • 在PushT数据集上,预测接口使决策成本降低0.002549(95%置信区间负值)
  • 适合关注推理效率与模型切换策略的强化学习研究者

现有自适应推理与世界-动作模型系统常利用廉价阶段输出或预测未来来分配额外计算。本文聚焦更窄问题:在配对精确重置的物理结果下,基于中等模型生成的接口能否准确预测切换到独立冻结的全模型,是否足以降低任务特定决策损失,从而证明串行开销合理?贡献在于提出一套配对评估与审计协议,而非新通用路由规则:所有候选动作均从相同重置状态执行,中等与全模型作用于同一候选集与任务,其配对物理损失差定义路由目标。在全新的PushT数据集(V106;1,600个状态,39个任务,三组检查点)上,冻结的预测接口路由器相比独立中等、独立全模型及延迟占优的任务专用路由器,降低了包含开销的决策成本。随后,前瞻性地对第二个1,600状态的PushT验证(V107)进行封存,采用任务信息、当前DINO特征的维度匹配投影及全部五个候选动作,且不计DINO编码器延迟。预测接口使定价后的物理决策成本降低0.002549(状态聚类95%区间[-0.002867, -0.002238];单侧95%上限-0.002286),对三组检查点均呈负向影响。受控的PyBullet审计独立支持复合任务-预测模式路由器。串行路由器仍慢于固定策略,其优势仅限于低计算成本场景。证据表明,在测试的预测接口中存在超出单一刻意偏好的当前DINO控制的增量路由信息,但不足以支撑因果充分性、计算节省、闭环价值或跨家族泛化。

原文摘要 · Abstract (English)

Existing adaptive-inference and world-action-model systems use cheap-stage outputs or predicted futures to allocate additional computation. We study a narrower question: under paired exact-reset physical outcomes, can a Medium-derived interface predict when switching to a separately frozen Full predictor improves task-specific decision loss enough to justify sequential overhead? Our contribution is a paired evaluation and audit protocol, not a new generic routing rule: all candidate actions are executed from the same reset state, Medium and Full act on the same candidate set and task, and their paired physical-loss difference defines the routing target. On a fresh PushT bank (V106; 1,600 states, 39 tasks, three checkpoint pairs), a frozen prediction-interface router lowers overhead-inclusive decision cost relative to standalone Medium, standalone Full, and a latency-advantaged task-only router. We then prospectively seal a second 1,600-state PushT confirmation (V107) against a stronger current-state control using the task, a dimension-matched projection of current DINO features, and all five candidate actions, with no DINO encoder latency charged. The prediction interface lowers priced physical decision cost by 0.002549 (state-clustered 95% interval [-0.002867, -0.002238]; one-sided 95% upper bound -0.002286), with negative effects for all three checkpoint pairs. A controlled-PyBullet audit independently supports a composite task-prediction-regime router. The sequential router remains slower than fixed policies, and its advantage is restricted to low compute prices. The evidence supports incremental routing information in the tested prediction interface beyond one deliberately favoured current-DINO control, but not causal sufficiency, compute saving, closed-loop value, or cross-family generality.

强化学习模型切换推理优化世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。