通过具身数据金字塔联合训练,提升视觉语言动作模型的物理理解能力。
CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understanding

- 构建具身物理VQA数据集与基准,对齐机器人动作域。
- 引入GAP令牌隔离运动规律,使动作头专注物理常识。
- 在真实和仿真任务中均显著优于基线模型。
视觉-语言-动作(VLA)模型在需要物理常识的操控任务中仍表现脆弱。现有物理视觉问答(VQA)数据多为非具身且与机器人动作域不匹配,而第一人称视频仅用作辅助预训练。尚不清楚更强的VLM物理理解是否真能提升下游动作生成性能。为此,我们提出CometVLA,构建CometData与CometBench——一个严格对齐机器人动作数据与具身性的具身物理VQA语料库与基准。提出全局动作先验(GAP)令牌,一种紧凑可学习瓶颈,分离任务无关的运动规律,使动作头能消费物理常识而不污染预训练的VLM主干。我们在涵盖远程操控、仿真、第一人称轨迹与VQA层的具身数据金字塔上联合训练CometVLA。在真实世界操控任务与RoboTwin仿真环境中,CometVLA持续优于强基线。相关性分析显示,VLM在CometBench上的性能越强,对应VLA成功率达越高。结果证明,物理理解预训练确实能有效提升下游操控性能。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Current physical VQA data is typically disembodied and misaligned with robot action domains. Egocentric videos are used only as auxiliary pre-training. It remains unclear whether improved VLM physical understanding actually benefits downstream action generation. Therefore, we present CometVLA to close this gap. We construct CometData and CometBench, an embodied physical VQA corpus and benchmark strictly aligned with the robot's action data and embodiment. We introduce Global Action Prior (GAP) tokens, a compact learnable bottleneck that isolates task-agnostic motion regularities and lets the action head consume physical commonsense without corrupting the pre-trained VLM backbone. We co-train CometVLA across the embodied data pyramid, spanning teleoperation, simulation, egocentric trajectories, and VQA layers. On real-world manipulation tasks and RoboTwin simulation, CometVLA consistently outperforms strong VLA baselines. Correlation analysis shows that stronger VLM performance on CometBench indicates higher VLA success rates. Results demonstrate that physical understanding pre-training genuinely benefits downstream manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。