arXiv:2509.14889cs.RO2025-09被引 5

CollabVLA让视觉语言动作模型能自我反思并主动求助,提升协作效率与可解释性。

CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human

  • 融合视觉语言模型推理与扩散动作生成,采用专家混合架构
  • 自反射机制使成功率更高,延迟降低约50%,梦境次数减少75%
  • 适合需要人机协同的复杂任务场景,如机器人操作与智能助手

本文提出CollabVLA,一种具备自反思能力的视觉-语言-动作框架,将标准视觉运动策略转化为协作型助手。该模型通过融合基于VLM的反思推理与基于扩散的动作生成,在专家混合架构下解决了先前模型存在的领域过拟合、推理不可解释及生成模型高延迟等问题。采用两阶段训练:动作对齐与反思调优,支持显式自我反思,并在不确定性或重复失败时主动寻求人类指导。相比生成式代理,其归一化时间减少约2倍,梦境次数减少约4倍,实现更高成功率、更好可解释性与更低延迟,推动视觉语言动作模型从黑箱控制向真正协作智能体演进。

原文摘要 · Abstract (English)

In this work, we present CollabVLA, a self-reflective vision-language-action framework that transforms a standard visuomotor policy into a collaborative assistant. CollabVLA tackles key limitations of prior VLAs, including domain overfitting, non-interpretable reasoning, and the high latency of auxiliary generative models, by integrating VLM-based reflective reasoning with diffusion-based action generation under a mixture-of-experts design. Through a two-stage training recipe of action grounding and reflection tuning, it supports explicit self-reflection and proactively solicits human guidance when confronted with uncertainty or repeated failure. It cuts normalized Time by ~2x and Dream counts by ~4x vs. generative agents, achieving higher success rates, improved interpretability, and balanced low latency compared with existing methods. This work takes a pioneering step toward shifting VLAs from opaque controllers to genuinely assistive agents capable of reasoning, acting, and collaborating with humans.

人机协作自反思视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。