让机器人通过具身思维链更智能地完成复杂操作,突破传统推理模式的局限。
Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation

- 用真实机器人数据构建最大规模具身思维链语料库,支持大规模训练
- 提出新模型ERVLA,训练时吸收推理过程,推理时不依赖思维链生成
- 在多个任务上实现领先性能,尤其擅长长序列和语义模糊任务
具身思维链(Embodied Chain-of-Thought)旨在连接语言推理与机器人控制,但其有效形式与融合策略仍不明确。本文大规模重构视觉-语言-动作(VLA)模型中的具身思维链,构建迄今最大的具身思维链语料库,包含978,743条轨迹、226.3M样本及2592.5小时机器人数据。实验发现,有效的具身思维链应将高层语义理解转化为具体动作指导,如末端执行器运动描述或图像空间轨迹,而仅依赖高层推理带来的收益微乎其微。此外,将显式思维链作为自回归动作前缀使用时,会因推理误差累积和推理-动作耦合不稳定而不可靠。为此,我们提出ERVLA,一种以具身思维链作为表征塑造监督而非测试时强制推理的VLA模型。该模型采用推理丢弃训练策略,在训练中吸收丰富推理痕迹,推理阶段直接预测动作,无需解码思维链。该设计提升了预训练数据增长下的可扩展性,并避免了自回归不稳定性。ERVLA在LIBERO-Plus上取得86.9%的成功率,在VLABench上达到53.2%的成功率,展现出强大的分布外泛化能力。真实机器人实验进一步表明,其优于现有先进基线,尤其在需语义消歧与长时序执行的任务中表现突出。
原文摘要 · Abstract (English)
Embodied chain-of-thought (CoT) aims to bridge linguistic reasoning and robotic control, but its effective form and integration strategy remain underexplored. In this paper, we revisit embodied CoT for vision-language-action (VLA) models at large scale. We construct the largest embodied CoT corpus to date, comprising 978,743 trajectories, 226.3M samples, and 2592.5 hours of robot data. Through extensive experiments, we find that effective embodied CoT should ground high-level semantic understanding into concrete action guidance, such as end-effector movement descriptions and image-space trajectories, while high-level reasoning alone brings only marginal gains. We further show that explicit CoT does not scale reliably when used as an autoregressive action prefix, as it suffers from compounding inference errors and unstable reasoning-action coupling. To address these limitations, we propose ERVLA, a VLA model that uses embodied CoT as representation-shaping supervision rather than mandatory test-time reasoning. ERVLA is trained with a reasoning-dropout strategy, enabling the model to absorb rich reasoning traces during training while predicting actions directly without CoT decoding during inference. This design improves scalability with increasing pre-training data and avoids autoregressive instability. ERVLA achieves state-of-the-art performance on LIBERO-Plus with an 86.9% success rate and reaches 53.2% success rate on VLABench, demonstrating strong out-of-distribution generalization. In real-robot experiments, ERVLA further outperforms competitive state-of-the-art baselines, especially on tasks requiring semantic disambiguation and long-horizon execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。