arXiv:2606.26801cs.ROcs.CV2026-06被引 2

通过自动识别操作阶段和关键帧目标,提升机器人长时序操作的训练效果。

Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision

论文配图:Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision
图 1 · 摘自论文原文
  • 引入阶段分类器与关键帧预测器,实现无需人工标注的结构化监督
  • 在真实机械臂任务中成功率提升56%,长任务效果更显著
  • 适配现有模型架构,可直接用于复杂操作任务的微调

视觉-语言-动作(VLA)模型在通用机器人操作中展现出巨大潜力。然而,在微调过程中,动作监督对所有时间步一视同仁,缺乏对当前操作阶段或下一夹持事件目标的结构化指导,导致失败集中在困难的夹持转换环节。为此,我们提出StaKe——一个无需人工标注即可从演示夹持状态自动提取互补信号的插件式辅助监督框架:包含阶段分类器以识别当前操作阶段,以及关键帧预测器以估计下一次夹持转换的目标关节动作。两者作为轻量级辅助头,丰富训练过程中的表征学习,同时保持原有VLA策略架构与推理流程不变。在双臂仿真和单臂Franka真实机器人任务上的实验表明,StaKe分别带来14%和56%的相对成功率提升,尤其在涉及更多夹持转换的长时序任务中优势更明显。消融实验验证了各项设计的有效性,定性分析也证实学习到的表征能准确追踪操作阶段。结果表明,结构化监督是提升长时序操作中VLA微调效果的有效且通用策略。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine-tuning, however, action supervision applies equally across all timesteps, without structured supervision on which manipulation stage the robot is in or what the next gripper-event target should be. This causes failures to concentrate around challenging gripper-event transitions. To address this, we propose StaKe, a plug-in auxiliary supervision framework that automatically derives two complementary signals from demonstration gripper states without manual annotation: a stage classifier that identifies the current manipulation stage, and a keyframe predictor that estimates the target joint action at the next gripper transition. Both are modeled as lightweight auxiliary heads that enrich the learned representations during training, while leaving the base VLA policy architecture and inference loop unchanged. Experiments on bimanual simulation and single-arm Franka real-robot tasks show that StaKe consistently improves success rates (relative gains of 14% and 56%, respectively), with larger improvements on longer-horizon tasks that involve more gripper-event transitions. Ablation studies validate each design choice, and qualitative analysis confirms that the learned representations faithfully track manipulation stages. These results indicate that structured supervision is an effective and general strategy for enhancing VLA fine-tuning in long-horizon manipulation. Project website: https://hi-yuanxu.github.io/StaKe-Web/

机器人操作结构化监督多模态学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。