对比视觉语言动作模型与世界动作模型的行为差异,发现未来预测能提升操作精细度。
Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA

- 通过行为轨迹与特征空间双视角诊断模型行为
- 序列式世界动作模型具最明显未来预测结构
- 适合关注机器人控制细节与模型可解释性的研究者
视觉-语言-动作(VLA)策略与世界-动作模型(WAM)是机器人操作中日益重要的范式。然而,现有研究未明确未来预测在WAM中是否带来超越任务成功率的行为意义。本文提出一种模型无关的诊断框架,从行为滚动分析和稀疏自编码器特征分析两个角度比较WAM与VLA。行为协议评估动作动态一致性、目标物体进展、干扰物扰动及运行成本;特征空间协议揭示内部表征是否为记忆型、反应型或预测型。在LIBERO与RoboTwin2.0数据集上测试7种策略,涵盖直接VLA及联合、序列、辅助型WAM。结果表明:仅看成功指标会掩盖关键差异——WAM普遍提升物体级行为与目标选择性,但性能依赖架构且推理开销更高。序列型WAM展现出最清晰的预测结构,而辅助与联合型分别压缩或纠缠未来信息。这些发现为设计保留可操作未来表征的高效操控模型指明方向。
原文摘要 · Abstract (English)
Vision-language-action (VLA) policies and World-Action Models (WAM) represent two increasingly important paradigms for robotic manipulation. However, it remains unclear whether future prediction in WAMs leads to behaviorally meaningful improvements beyond final task success. In this paper, we ask whether WAMs merely add future prediction, or whether they change robot behavior and internal representations in ways that are actionable for control. We introduce a model-agnostic diagnostic framework that compares WAMs and VLAs through two complementary lenses: behavioral rollout analysis and sparse-autoencoder-based feature analysis. The behavioral protocol measures action dynamics consistency, target-object progress, distractor disturbance, and runtime cost. The feature-space protocol characterizes internal representations as memorized, reactive, or predictive, revealing whether models encode future-oriented structure. Across LIBERO and RoboTwin2.0, we evaluate 7 policies spanning direct VLAs and joint, sequential, and auxiliary WAMs. Our results show that success alone hides key differences: WAMs often improve object-level behavior and target selectivity, but their gains depend on architecture and incur higher inference cost. Sequential WAMs show the clearest predictive structure, while auxiliary and joint WAMs respectively compress or entangle future information. These findings suggest future directions for WAMs design to preserve behaviorally actionable future representations for efficient manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。