用可表述的潜在思维实现高效视觉语言动作推理,速度提升近90%。
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
- 通过压缩的隐式思维链替代长推理过程,提升效率。
- 在多种任务中实现89.3%的推理延迟降低,同时保持长程规划能力。
- 适合需要快速响应的机器人控制与少样本适应场景。
视觉-语言-动作(VLA)任务要求对复杂视觉场景进行推理,并在动态环境中执行自适应动作。尽管近期研究显示显式思维链(CoT)能提升泛化能力,但其推理轨迹过长导致推理延迟高。我们提出 Fast-ThinkAct,一种高效的推理框架,通过可表述的潜在思维链实现紧凑而高性能的规划。该方法通过教师模型蒸馏学习,采用偏好引导目标,对齐操作轨迹,同时传递语言与视觉规划能力以支持具身控制。这使得增强推理的策略学习能够有效连接紧凑推理与动作执行。在多个具身操作与推理基准上的广泛实验表明,Fast-ThinkAct 在保持长程规划、少样本适应和失败恢复能力的同时,相较当前最优推理 VLA 推理延迟最高降低 89.3%。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) tasks require reasoning over complex visual scenes and executing adaptive actions in dynamic environments. While recent studies on reasoning VLAs show that explicit chain-of-thought (CoT) can improve generalization, they suffer from high inference latency due to lengthy reasoning traces. We propose Fast-ThinkAct, an efficient reasoning framework that achieves compact yet performant planning through verbalizable latent reasoning. Fast-ThinkAct learns to reason efficiently with latent CoTs by distilling from a teacher, driven by a preference-guided objective to align manipulation trajectories that transfers both linguistic and visual planning capabilities for embodied control. This enables reasoning-enhanced policy learning that effectively connects compact reasoning to action execution. Extensive experiments across diverse embodied manipulation and reasoning benchmarks demonstrate that Fast-ThinkAct achieves strong performance with up to 89.3% reduced inference latency over state-of-the-art reasoning VLAs, while maintaining effective long-horizon planning, few-shot adaptation, and failure recovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。