让AI像人一样思考再行动,实现复杂任务的自主规划与纠错
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

- 用视觉隐变量规划来连接思考与执行,分两步完成任务
- 在多个机器人任务上实现少样本适应和长达10步的计划能力
- 适合研究具身智能、机器人控制或多模态推理的开发者
视觉-语言-动作(VLA)推理任务要求智能体理解多模态指令,在动态环境中进行长程规划并自适应执行。现有方法通常采用端到端训练,直接从输入映射到动作,缺乏显式的推理过程,难以实现多步规划或应对复杂任务变化。本文提出ThinkAct,一种双系统框架,通过强化的视觉隐变量规划,将高层推理与底层动作执行相连接。该框架训练多模态大模型生成受动作对齐视觉奖励(基于目标达成和轨迹一致性)引导的具身推理计划,再将这些计划压缩为视觉计划隐变量,用于指导下游动作模型在目标环境中的鲁棒执行。在具身推理与机器人操作基准上的大量实验表明,ThinkAct可实现少样本适应、长程规划(最高达10步)及自我纠错行为,在复杂具身人工智能任务中表现优异。
原文摘要 · Abstract (English)
Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end fashion, directly mapping inputs to actions without explicit reasoning, which hinders their ability to plan over multiple steps or adapt to complex task variations. In this paper, we propose ThinkAct, a dual-system framework that bridges high-level reasoning with low-level action execution via reinforced visual latent planning. ThinkAct trains a multimodal LLM to generate embodied reasoning plans guided by reinforcing action-aligned visual rewards based on goal completion and trajectory consistency. These reasoning plans are compressed into a visual plan latent that conditions a downstream action model for robust action execution on target environments. Extensive experiments on embodied reasoning and robot manipulation benchmarks demonstrate that ThinkAct enables few-shot adaptation, long-horizon planning, and self-correction behaviors in complex embodied AI tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。