arXiv:2608.23863cs.RO2026-08

用可追溯的信用记录机制,让机器人拒绝不可靠的预测。

DreamLedger: Where to Refuse World-Model Imagination Using Execution-Settled Credit

论文配图:DreamLedger: Where to Refuse World-Model Imagination Using Execution-Settled Credit
图 1 · 摘自论文原文
  • 将预测可靠性存为持久信用档案,按环境、区域和时间窗口记录
  • 实测减少62%的无效想象,故障预测命中率提升69%
  • 适合需要高可靠性决策的机器人系统与真实硬件部署

机器人开始基于世界模型进行预测行动,但当前可靠性仍依赖瞬时的内部信号,无法反映过往想象失败的历史。本文提出DreamLedger,将可靠性视为持续可追踪的部署对象:一个执行后结算的信用档案,记录每条预测在不同运行条件、区域和预测时长下的兑现情况。每个预测被消费后自动注册为索赔,并在现实反馈到来时结算,无需人工标注。信用值决定是否使用该预测(低信用则缩短依赖或触发观察),所有依赖事件可通过依赖票据和可重放日志审计。信用变化反映的是‘拒绝’行为而非模型出错:69%的拒绝发生在已失败的区域;在健康条件下,局部重置使错误拒绝翻倍;局部递归退化下,信用机制使消耗的想象减少一半,任务完成率略有下降。在三个模拟域及真实Franka机械臂和双旋翼无人机上测试,未修改的DreamerV3、TD-MPC2、V-JEPA 2-AC模型均显示,信用门控规划相比盲目使用减少62%(95% CI 43-81%)的无效想象;基于结算校准的策略在种子一致的中等可信度点运行,而原始瞬时阈值则陷入极端。信用层覆盖解码器、隐空间与标记空间接口。在硬件上,结算机制在真实传感与接触噪声下运行,两模型均在9厘米冻结容忍度下被评定为可信,故障循环在线重新定价,全部1,062次使用记录均可从审计日志回放。

原文摘要 · Abstract (English)

Robots are beginning to act on world-model predictions, yet reliability is still expressed through instantaneous, model-internal signals that say whether a prediction looks trustworthy now, not where comparable imagination has already failed. DreamLedger instead treats reliability as a persistent deployment object: an execution-settled credit file recording how often consumed predictions are borne out, indexed by operating condition, region, and prediction horizon, and consulted before each use. Each consumed prediction is registered as a claim and settled against arriving reality without manual labels; the resulting credit gates consumption (low credit shortens reliance or triggers observation), and every reliance event remains auditable via dependency tickets and replayable logs. Persistent credit changes where the gate refuses rather than what the model gets wrong: 69% of denials land on cells that have already failed, episode-local reset triples off-target denials in healthy conditions, and under a localized recurrent degradation persistent credit halves burned imagination, at a cost in task completion. Across three simulated domains, unmodified DreamerV3, TD-MPC2, and V-JEPA 2-AC mounts, and a real Franka, paired quadrotor evaluation shows credit-gated planning reduces burned imagination by 62% (95% CI 43-81%) versus blind consumption; settlement-grounded calibration yields moderate, seed-consistent operating points where raw instantaneous gates collapse to extremes, while persistent books trade verification for reliance (manipulation probes 1.00 to 0.36/episode at success 0.98 vs. 0.94). The trust layer spans decoder-, latent-, and token-space interfaces. On hardware, settlement runs under real sensing and contact noise, both models are priced creditworthy at the frozen 9-cm tolerance, a failure loop is re-priced online, and all 1,062 registered spends replay from the audit logs.

机器人决策可信推理信用机制真实部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。