让视觉语言动作模型提前预判未来动态,提升机器人操作成功率。
PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models

- 用未来观测提取的隐状态引导当前决策,增强前瞻性推理。
- 在LIBERO数据集上成功率从84.1%提升至88.4%,真实拆解任务从63.3%到82.5%。
- 适合需要精细接触控制的机器人操作任务,尤其关注未来规划的场景。
视觉-语言-动作模型(VLAs)通过直接将语言指令和视觉观测映射为动作,在通用机器人操作中展现出巨大潜力。然而,大多数VLAs仅依赖当前观测进行动作预测,缺乏对未来的任务动态进行显式推理的机制,这在精细、高接触密度的操作任务中尤为关键。本文提出PHR-VLA框架,通过引入未来动态的特权隐表示,实现VLAs中的规划时域推理。该框架在训练阶段加入一个轻量级辅助未来头,使VLA的内部表示与来自未来观测的隐动态对齐。实验表明,从手腕摄像头获得的局部、接触中心的像素级隐动态监督,使LIBERO上的成功率从84.1%提升至88.4%,真实世界拆卸任务从63.3%提升至82.5%;第三人称视角的像素级监督也使Meta-World上的性能从56.70%提升至57.8%。结果证明,特权隐动态对齐为提升VLA策略的预见性推理提供了有效训练信号。
原文摘要 · Abstract (English)
Vision-language-action models (VLAs) have shown strong promise for general-purpose robotic manipulation by mapping language instructions and vision observations directly to actions. However, most VLAs primarily condition action prediction on current observations and lack an explicit mechanism for reasoning over future task dynamics, which is particularly important for fine-grained, contact-rich manipulation. We present PHR-VLA, a framework that enables planning-horizon reasoning in VLAs through privileged latent representations of future dynamics. PHR-VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations. Evaluation results demonstrate that local, contact-centric, patch-level latent dynamics supervision from the wrist camera improves success rate on LIBERO from 84.1% to 88.4% and on real-world disassembly tasks from 63.3% to 82.5%. Patch-level supervision from a third-person camera also improves performance on Meta-World from 56.70% to 57.8%. These results demonstrate that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA policies. Project website: \href{https://davoodsz.github.io/PHR-VLA.github.io/}{https://davoodsz.github.io/PHR-VLA.github.io/}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。