arXiv:2606.06245cs.ROcs.AI2026-06

让视觉语言动作模型在推理时多路思考,不增加延迟却提升长任务表现。

MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action

论文配图:MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action
图 1 · 摘自论文原文
  • 用多路径潜空间推理替代传统链式思考,避免生成额外文本
  • 在LIBERO和CALVIN上实现8步动作接口下的性能提升
  • 支持可调推理深度与宽度,适合高不确定性长程控制场景

视觉-语言-动作(VLA)策略在长周期和高不确定性控制中仍显脆弱,单次动作解码难以实现充分的推理。虽然显式思维链能加深推理,但会引入词元延迟并使用间接的文本到动作接口。我们提出MPCoT,一种奖励引导的多路径潜空间推理框架:初始化M条假设路径,经K次共享权重的迭代优化后,通过软聚合生成动作。训练阶段采用路径偏好目标,融合专家轨迹一致性、冻结的Qwen3-VL进展评分及终点成功反馈,使潜空间路径评分与实际执行质量对齐。MPCoT保持原始8步动作接口,生成零推理词元,并支持可配置的推理控制参数(K, M)。在LIBERO和CALVIN上匹配协议测试中,该方法显著提升长周期任务性能;消融实验证明了推理深度与宽度的影响、置信度加权聚合以及奖励引导路径监督的有效性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) policies remain brittle in long-horizon and high-uncertainty control, where one-pass action decoding provides limited inference-time deliberation. Explicit chain-of-thought can increase reasoning depth, but introduces token latency and an indirect text-to-action interface. We propose MPCoT, a reward-guided multi-path latent reasoning framework that initializes M hypotheses, refines them for K weight-tied steps, and softly aggregates them before action decoding. A training-only path-preference objective combines expert-trajectory consistency, frozen Qwen3-VL progress scoring, and endpoint-success feedback to align the latent path scorer with downstream execution quality. MPCoT preserves the original 8-step action interface, generates zero reasoning tokens, and exposes configurable inference controls (K, M). Under matched protocols on LIBERO and CALVIN, MPCoT improves long-horizon performance, with ablations confirming depth-width effects, confidence-weighted aggregation, and reward-guided path supervision.

视觉语言动作多路径推理奖励引导长程控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。