arXiv:2601.22467cs.ROcs.CV2026-01中稿 · 2026 IEEE Internat…被引 1

用视频和文字训练机器人,不用动作标注也能学会连续控制动作。

CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control

  • 用视频-文本对预训练,不依赖动作标签学习连续动作表示。
  • 在仿真任务中成功率显著提升,避免了捷径学习问题。
  • 适合做弱监督机器人控制,尤其适用于数据标注难的场景。

视觉-语言-动作(VLA)模型在机器人控制中展现出潜力,但其对动作标注的依赖限制了可扩展性和泛化能力。为解决此问题,我们提出CARE框架,通过仅使用视频-文本对进行预训练,无需显式动作标签即可学习连续的潜在动作表示。该方法采用新设计的多任务预训练目标,使模型在微调阶段仅需少量标注数据即可训练动作头完成控制。实验结果表明,CARE在多个仿真任务中表现出更高的成功率、更好的语义可解释性,并能有效避免捷径学习。这些成果证明了CARE在弱监督机器人控制中的可扩展性、可解释性与有效性。

原文摘要 · Abstract (English)

Recent advances in Vision-Language-Action (VLA) models have shown promise for robot control, but their dependence on action supervision limits scalability and generalization. To address this challenge, we introduce CARE, a novel framework designed to train VLA models for robotic task execution. Unlike existing methods that depend on action annotations during pretraining, CARE eliminates the need for explicit action labels by leveraging only video-text pairs. These weakly aligned data sources enable the model to learn continuous latent action representations through a newly designed multi-task pretraining objective. During fine-tuning, a small set of labeled data is used to train the action head for control. Experimental results across various simulation tasks demonstrate CARE's superior success rate, semantic interpretability, and ability to avoid shortcut learning. These results underscore CARE's scalability, interpretability, and effectiveness in robotic control with weak supervision.

机器人控制弱监督多任务预训练连续动作表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。