arXiv:2602.12532cs.RO2026-02被引 6

让机器人更精准地完成需要接触力的复杂操作。

CRAFT: Adapting VLA Models to Contact-rich Manipulation via Force-aware Curriculum Fine-tuning

  • 通过力觉感知引导训练,逐步融合视觉语言信息
  • 在真实场景中提升接触密集任务成功率,泛化能力更强
  • 适合做机器人灵巧操作与多模态学习的研究者

视觉-语言-动作(VLA)模型在执行通用指令方面表现出色,但在需要精确对齐、稳定接触和处理柔顺物体的接触密集型操作任务中表现不佳。根本问题在于高熵的视觉与语言输入与低熵但关键的力觉信号之间存在不平衡,导致模型过度依赖感知而控制不稳。为此,我们提出CRAFT,一种基于力觉感知的课程微调框架,引入变分信息瓶颈模块,在训练初期调控视觉与语言嵌入,引导模型优先关注力觉信号,再逐步恢复全模态信息。为支持力觉学习,我们设计了同源主从遥操作系统,采集多样化接触任务下的同步视觉、语言与力觉数据。实验证明,CRAFT能持续提升任务成功率,有效泛化至未见物体与新任务变化,并适配多种VLA架构,实现鲁棒且通用的接触密集型操作能力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown a strong capability in enabling robots to execute general instructions, yet they struggle with contact-rich manipulation tasks, where success requires precise alignment, stable contact maintenance, and effective handling of deformable objects. A fundamental challenge arises from the imbalance between high-entropy vision and language inputs and low-entropy but critical force signals, which often leads to over-reliance on perception and unstable control. To address this, we introduce CRAFT, a force-aware curriculum fine-tuning framework that integrates a variational information bottleneck module to regulate vision and language embeddings during early training. This curriculum strategy encourages the model to prioritize force signals initially, before progressively restoring access to the full multimodal information. To enable force-aware learning, we further design a homologous leader-follower teleoperation system that collects synchronized vision, language, and force data across diverse contact-rich tasks. Real-world experiments demonstrate that CRAFT consistently improves task success, generalizes to unseen objects and novel task variations, and adapts effectively across diverse VLA architectures, enabling robust and generalizable contact-rich manipulation.

机器人操作多模态学习力觉控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。