arXiv:2602.02459cs.RO2026-02被引 11

让机器人在动态环境中听懂指令并实时反应,解决语义推理延迟问题。

TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments

  • 引入延迟感知的控制框架,将语义推理延迟显式纳入决策
  • 在模拟和真实机器人上均实现多秒延迟下的稳定导航
  • 适合需要实时响应的智能机器人应用,如服务机器人

动态人机交互环境中,机器人需遵循语言指令并保持实时反应能力。视觉-语言-动作(VLA)模型虽具潜力,但假设语义推理与控制同步,而实际中语义理解存在固有延迟。本文提出延迟感知的思考-控制(TIC-VLA)框架,显式建模推理延迟,通过延迟的视觉-语言语义状态和延迟元数据,结合当前观测,指导动作生成,使策略能补偿异步推理。进一步设计一致性延迟训练流程,在模仿学习和在线强化学习中注入推理延迟,使训练与部署对齐。为支持真实评估,构建了DynaNav——一个物理精确、照片级真实的动态环境语言导航仿真套件。大量实验表明,TIC-VLA在仿真与真实机器人上均显著优于现有VLA模型,且在多秒推理延迟下仍保持稳健实时控制。

原文摘要 · Abstract (English)

Robots in dynamic, human-centric environments must follow language instructions while maintaining real-time reactive control. Vision-language-action (VLA) models offer a promising framework, but they assume temporally aligned reasoning and control, despite semantic inference being inherently delayed relative to real-time action. We introduce Think-in-Control (TIC)-VLA, a latency-aware framework that explicitly models delayed semantic reasoning during action generation. TIC-VLA defines a delayed semantic-control interface that conditions action generation on delayed vision-language semantic states and explicit latency metadata, in addition to current observations, enabling policies to compensate for asynchronous reasoning. We further propose a latency-consistent training pipeline that injects reasoning inference delays during imitation learning and online reinforcement learning, aligning training with asynchronous deployment. To support realistic evaluation, we present DynaNav, a physics-accurate, photo-realistic simulation suite for language-guided navigation in dynamic environments. Extensive experiments in simulation and on a real robot show that TIC-VLA consistently outperforms prior VLA models while maintaining robust real-time control under multi-second reasoning latency. Project website: https://ucla-mobility.github.io/TIC-VLA/

机器人导航视觉语言动作延迟感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。