让机器人决策更可信:提升视觉语言动作模型的思维真实性
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning

- 用可学习的评鉴器Pinocchio量化思维与环境的贴合度和逻辑连贯性
- 后训练使思维真实性提升4%~18%,任务表现仍保持领先
- 适合关注机器人可解释性与鲁棒性的研究者
具身链式思考(Embodied Chain-of-Thought)被视作提升视觉语言动作(VLA)模型决策能力与可解释性的关键机制。然而,其生成的推理是否真实反映策略内部决策过程尚不明确。本文区分功能性推理(提升任务性能)与真实性推理(真实反映内部决策)。现有顶尖对齐策略虽必要但不足,可能掩盖动作预测中的因果关联(如脱离环境、逻辑断裂)。通过人类评估发现,顶尖推理模型的推理质量与轨迹改进存在不一致。为此,提出可学习评鉴器Pinocchio,以观察接地性与步骤连贯性为代理指标,作为强化学习中的密集奖励信号进行后训练。在保留下游任务性能的前提下,新规划器在多个驾驶基准上分别较顶尖对齐与轨迹误差基线提升4%与18%的真实思维比例。合成分布外测试集显示,真实思维后训练使政策对罕见反事实场景响应能力提高1.6倍,表明真实性推理有助于构建更鲁棒、可泛化、可解释的智能体。
原文摘要 · Abstract (English)
Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-Language Action (VLA) models. However, whether this verbalized Chain-of-Thought truthfully reflects the policy's underlying decision process remains poorly understood. We distinguish between functional reasoning, in which reasoning improves task performance, and faithful reasoning, in which reasoning truly reflects the policy's internal decision process. We argue that SoTA alignment strategies offer a necessary but insufficient notion of faithfulness, admitting reasoning whose intermediate steps can mask the causal links in action prediction through confounding factors (e.g., reasoning that is ungrounded in the environment and internally disconnected or inconsistent), restricting policy generalization. We study this gap through a human evaluation of a SoTA reasoning model for autonomous driving, revealing an inconsistent coupling between reasoning quality and downstream trajectory improvement. We then operationalize a behavioral surrogate for embodied faithfulness through a learned critic, Pinocchio, scoring observation grounding and stepwise coherence, and use this critic as a dense reward signal in post-training an embodied policy with reinforcement learning. Across withheld driving benchmarks, our post-trained planner improves faithfulness by 4% and 18% over SoTA alignment and trajectory error post-training baselines, respectively, while maintaining competitive downstream task performance. Finally, on a synthetic out-of-distribution test set, post-training for faithfulness improves policy responsiveness to rare counterfactual scenarios by 1.6x that of a SoTA policy, suggesting that faithful reasoning traces contribute to more robust, generalizable, and interpretable embodied intelligence. Project page: https://mjf-su.github.io/pinocchio/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。