arXiv:2605.05126cs.RO2026-05中稿 · CVPR被引 3

提升机器人操作中的时空一致性,实现高效3D感知与4D推理。

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation

论文配图:ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
图 1 · 摘自论文原文
  • 通过多视角对齐与几何融合,增强物体语义与空间关系的一致性。
  • 在LIBERO基准上性能提升21.6%,推理速度加快2.4倍。
  • 适合需要精准动态场景理解的机器人任务,如复杂操作与交互。

当前视觉-语言-动作(VLA)模型主要将2D观测映射为动作,但在时空感知与推理方面存在明显局限:1)空间表征依赖额外传感器,带来显著计算开销;2)视觉推理通常仅限于未来帧预测,缺乏与指令引导场景的对齐,削弱了时空一致性。为此,我们提出ConsisVLA-4D,一个统一且高效的框架,以增强3D感知与4D推理中的时空一致性。具体设计包括:1)CV-Aligner,通过筛选指令相关区域并跨视角对齐物体身份,保证跨视图语义一致性;2)CO-Fuser,利用紧凑潜在表示消除不同视角间物体空间关系的歧义,确保跨物体几何一致性。在此基础上,引入CS-Thinker,实现动作执行过程中跨场景的时空一致性。其从CV-Aligner的物体语义标记和CO-Fuser的全局深度标记中学习局部动态隐知识,从而在场景变化下实现高效视觉推理。大量实验表明,得益于其高效的时空一致性设计,ConsisVLA-4D在LIBERO基准和真实平台上的性能分别提升21.6%和41.5%,推理速度分别加快2.3倍和2.4倍。

原文摘要 · Abstract (English)

Current Vision-Language-Action (VLA) models primarily focus on mapping 2D observations to actions, but exhibit notable limitations in spatiotemporal perception and reasoning: 1) spatial representations often rely on additional sensors, introducing substantial computational overhead; 2) visual reasoning is typically limited to future-frame prediction, lacking alignment with the instruction-grounded scene and thus compromising spatiotemporal consistency. To address these challenges, we propose ConsisVLA-4D, a unified and efficient framework that enhances spatiotemporal consistency in 3D perception and 4D reasoning. Specifically, we design: 1) CV-Aligner, which ensures cross-view object semantic consistency by filtering instruction-relevant regions and aligning object identities across multiple viewpoints; 2) CO-Fuser, which guarantees cross-object spatial geometric consistency by eliminating spatial relation ambiguities between objects across views using compact latent representations. Building upon these, we introduce 3) CS-Thinker to achieve cross-scene spatiotemporal consistency as actions unfold. It learns implicit knowledge of local dynamics from object-semantic tokens of CV-Aligner and global depth from geometric tokens of CO-Fuser, thereby enhancing efficient visual reasoning under scene variations. Extensive experiments demonstrate that, benefiting from its efficient spatiotemporal consistency design, ConsisVLA-4D achieves 21.6% and 41.5% performance improvements, along with 2.3-fold and 2.4-fold inference speedups compared to OpenVLA on the LIBERO benchmark and real-world platforms, respectively.ConsisVLA-4D is open-sourced and publicly available at

机器人操作时空一致性3D感知多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。