arXiv:2608.11174cs.RO2026-08

提出VIScore指标,精准评估世界模型规划能力。

VIScore: Diagnosing Planning-Relevant Quality in Latent World Models

论文配图:VIScore: Diagnosing Planning-Relevant Quality in Latent World Models
图 1 · 摘自论文原文
  • 设计VIScore,综合评估编码器、预测器与规划器的性能
  • 在跨任务成功率上相关性超0.75,优于现有指标
  • 适合研究世界模型设计与诊断的科研人员

将潜在空间调节为各向同性的高斯分布可为世界模型规划提供稳定且信息量最大的环境。然而,潜在空间特性与规划成功之间缺乏关联。我们通过比较SIGReg和VISReg两种具有相同分布目标但不同特性的正则化损失函数进行研究。相比SIGReg,VISReg在控制中心、尺度和形状正则化权重方面更具灵活性,更大的批量大小能实现更精细的分布逼近。尽管SIGReg在自监督学习中表现更好,却无助于规划;而VISReg在域外(OOD)数据集上提升了规划成功率。这促使我们深入理解影响规划成功率的关键因素。不同于仅关注编码潜在变量的现有指标,我们提出验证性-影响力-节制性评分(VIScore),量化给定编码特征下预测器的可达性与容量,以及基于搜索的规划器的幻觉程度。相较于直线性、物理状态探测和赋能度,VIScore在覆盖编码器、预测器与规划器的测量中,对规划成功率的解释能力更强,表现为显著的斯皮尔曼相关性。具体而言,VIScore在已见与未见模型及数据集上的跨任务成功率池中,始终达到超过0.75的斯皮尔曼相关性。此外,它是唯一在所有测试场景中校准误差低于常数拟合的指标,凸显了这三个维度在规划成功中的重要性。我们希望该指标能助力未来世界模型的设计与诊断研究。

原文摘要 · Abstract (English)

Regulating the latent space to an isotropic Gaussian distribution provides a stable and information-maximized landscape for world model planning. However, the latent space property and successful planning remain disconnected. We first study this by comparing SIGReg and VISReg, two regularization loss functions with the same distribution target but different properties. Compared with SIGReg, VISReg has more flexibility in controlling the weights of center, scale, and shape regularization, and a larger batch size brings a finer distribution approximation. We find that the former, despite being beneficial in self-supervised learning (SSL), does not help the planning, whereas the latter improves the planning success on out-of-domain (OOD) datasets. This motivates a deep understanding of the factors that correlate with the success rate. Unlike the previous metrics focusing on the encoded latent only, we propose the Veracity-Influence-Sobriety score (VIScore), a metric that quantifies the reachability and capacity of a predictor given the encoded feature, and the hallucination of the searching-based planner. Compared with straightness, physical-state probing, and empowerment, we show that, with the measurement covering encoder, predictor, and planner, VIScore explains the success rate better than the others, as reflected by a strong Spearman correlation. Specifically, VIScore consistently achieves a Spearman correlation over 0.75 on both seen and unseen models and datasets on the cross-task success rate pool. Moreover, VIScore is the only metric that has a calibration error below the constant fit across all testing scenarios, showcasing the importance of these three aspects in planning success. We hope this metric can help future studies on world model design and diagnosis.

世界模型规划评估潜在空间指标设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。