通过可解释性分析,实现对世界动作模型的无训练可控增强。
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

- 用机制可解释性分析激活空间中鲁棒性特征的线性可分性
- 提出WA-LQR控制器,在新任务上提升抗扰能力
- 适用于需强鲁棒性的机器人控制场景
世界动作模型(WAMs)虽能实现语义与物理感知控制,但在分布偏移下表现脆弱。本文利用机制可解释性研究鲁棒性相关扰动在WAM激活空间中的表征。对比成功与失败轨迹的激活模式,发现部分架构中关键特征呈现低维线性可分,而另一些则不然。由此启发,提出基于对比激活方向的无训练模型调优方法。同时,发现激活动态具有局部线性特性,可高效使用基于模型的最优控制实现反馈调节,构建出最小侵入性的降阶LQR控制器——世界动作线性二次调节器(WA-LQR)。机制评估预测Cosmos-Policy和DiT4DiT模型具备强可调性,而LingBot-VA较弱,与实际干预结果一致。在Cosmos-Policy和DiT4DiT上,WA-LQR将对比方向泛化至新任务,显著提升对相机、夹爪及视觉噪声扰动的鲁棒性,优于未调优和提示调优基线。
原文摘要 · Abstract (English)
World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This motivates the use of contrastive activation directions for training-free WAM steering. We also show that local linearity in WAM activation dynamics enables efficient feedback steering via model-based optimal control, yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller. Via mechanistic evaluations, we predict strong steerability in the Cosmos-Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations over unsteered and prompt steering baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。