arXiv:2510.21817cs.ROcs.CL2025-10

让机器人同时看听说动,还能实时响应打断,像人一样干活。

VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting

  • 双模型并行架构,实现视觉、语音、动作并发处理
  • 紧急停止和语音打断成功率极高,支持边说话边执行任务
  • 适合需要自然交互的机器人助手、人机协作场景

当前视觉-语言-动作(VLA)模型受限于僵化的静态交互模式,难以同时感知视觉、听觉、生成语言并执行动作,也无法动态应对用户实时打断,导致人机协作生硬且响应迟缓。为此,我们提出VITA-E,一种新型具身交互框架,支持行为并发与近实时中断。核心为双模型架构,两个并行的VLA实例分别作为“主动模型”与“待命模型”,使智能体可同时观察环境、倾听用户语音、生成回应并执行动作,具备类人多任务能力。我们进一步提出“模型即控制器”范式,通过微调视觉语言模型生成特殊指令标记,直接驱动系统行为,实现推理与控制的耦合。在物理人形平台上实验表明,VITA-E能可靠应对复杂交互场景,兼容多种双系统VLA模型,在紧急停机和语音打断任务中表现优异,成功实现语音与动作的并发执行。这为更自然、更强大的具身助手迈出了重要一步。

原文摘要 · Abstract (English)

Current Vision-Language-Action (VLA) models are often constrained by a rigid, static interaction paradigm, which lacks the ability to see, hear, speak, and act concurrently as well as handle real-time user interruptions dynamically. This hinders seamless embodied collaboration, resulting in an inflexible and unresponsive user experience. To address these limitations, we introduce VITA-E, a novel embodied interaction framework designed for both behavioral concurrency and nearly real-time interruption. The core of our approach is a dual-model architecture where two parallel VLA instances operate as an ``Active Model'' and a ``Standby Model'', allowing the embodied agent to observe its environment, listen to user speech, provide verbal responses, and execute actions, all concurrently and interruptibly, mimicking human-like multitasking capabilities. We further propose a ``model-as-controller'' paradigm, where we fine-tune the VLM to generate special tokens that serve as direct system-level commands, coupling the model's reasoning with the system's behavior. Experiments conducted on a physical humanoid platform demonstrate that VITA-E can reliably handle complex interactive scenarios. Our framework is compatible with various dual-system VLA models, achieving an extremely high success rate on emergency stops and speech interruptions while also successfully performing concurrent speech and action. This represents a significant step towards more natural and capable embodied assistants.

具身智能多模态交互人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。