arXiv:2506.22242cs.CV2025-06NeurIPS被引 61

用4D信息对齐机器人与场景坐标,提升视觉-语言-动作模型的时空推理能力。

4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration

  • 输入序列RGB-D图像,融合深度与时间信息以对齐坐标系。
  • 在仿真和真实场景中成功率达78.3%,超越OpenVLA约15个百分点。
  • 适合需要强空间感知与跨场景泛化的机器人预训练任务。

利用多样化机器人数据进行预训练仍面临挑战。现有方法通常以简单观测作为输入建模动作分布,但这些输入常不完整,导致条件动作分布分散,我们称之为坐标系混乱与状态混乱,严重影响预训练效率。为此,我们提出4D-VLA,通过在输入中引入4D信息有效缓解此类混乱。模型使用序列化RGB-D输入,将深度与时间信息融入视觉特征,对齐机器人与场景的坐标系统,赋予模型强大的时空推理能力,同时保持低训练开销。此外,我们设计记忆库采样策略,从历史图像中提取有信息量的帧,进一步提升效果与效率。实验表明,该预训练方法及架构组件显著提升模型性能。在模拟与真实世界实验中,模型成功率相较OpenVLA提升约15个百分点。为评估空间感知与新视角泛化能力,我们构建了多视角仿真基准MV-Bench。结果表明,本模型持续领先,展现出更强的空间理解与适应性。

原文摘要 · Abstract (English)

Leveraging diverse robotic data for pretraining remains a critical challenge. Existing methods typically model the dataset's action distribution using simple observations as inputs. However, these inputs are often incomplete, resulting in a dispersed conditional action distribution-an issue we refer to as coordinate system chaos and state chaos. This inconsistency significantly hampers pretraining efficiency. To address this, we propose 4D-VLA, a novel approach that effectively integrates 4D information into the input to mitigate these sources of chaos. Our model introduces depth and temporal information into visual features with sequential RGB-D inputs, aligning the coordinate systems of the robot and the scene. This alignment endows the model with strong spatiotemporal reasoning capabilities while minimizing training overhead. Additionally, we introduce memory bank sampling, a frame sampling strategy designed to extract informative frames from historical images, further improving effectiveness and efficiency. Experimental results demonstrate that our pretraining method and architectural components substantially enhance model performance. In both simulated and real-world experiments, our model achieves a significant increase in success rate over OpenVLA. To further assess spatial perception and generalization to novel views, we introduce MV-Bench, a multi-view simulation benchmark. Our model consistently outperforms existing methods, demonstrating stronger spatial understanding and adaptability.

机器人4D建模预训练多视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。