arXiv:2607.13605cs.RO2026-07

对比了三种阶段信息接口对机器人长程操作的微调效果

An Empirical Study on Stage-Information Interfaces for VLA Fine-Tuning

论文配图:An Empirical Study on Stage-Information Interfaces for VLA Fine-Tuning
图 1 · 摘自论文原文
  • 用阶段信息作为任务指令与动作间的中间表示
  • 序号型阶段状态在延续微调中表现最佳,成功率53.75%
  • 显式阶段信息未必提升性能,效果依赖训练方式

一个长程操作的高层指令可能涵盖多个动作阶段。本文采用分段动作标注作为任务指令与机器人动作块之间的中间表示。通过进度模块追踪当前阶段,动作策略接收阶段信息的方式包括当前阶段文本或归一化的序号阶段索引。我们在LIBERO-10数据集上,以GR00T N1.6为基础,在直接微调和从全任务指令基线继续微调两种设置下比较这些接口。直接微调下,全任务指令、当前阶段文本、序号阶段状态的平均成功率分别为57.45%、50.24%、54.36%,表明显式阶段信息并未自动提升策略性能。延续微调下,对应平均值为49.07%、50.00%、53.75%,序号阶段状态在全部三组配对实验中均优于其他两者,且其增益效果随接口形式和训练方式变化而异。

原文摘要 · Abstract (English)

One high-level instruction in long-horizon manipulation can cover several action stages. We use segmented action annotations as an intermediate representation between the full-task instruction and VLA action chunks. A progress module tracks the active stage, while the action policy receives stage information either as current-stage text or as a normalized ordinal stage index in robot state. We compare these interfaces with GR00T N1.6 on LIBERO-10 under direct fine-tuning and continuation fine-tuning from a full-task instruction baseline. Under direct fine-tuning, full-task instruction, current-stage text, and Ordinal Stage-State achieve mean success rates of 57.45%, 50.24%, and 54.36%, respectively, showing that explicit stage information does not automatically improve the policy. Under continuation, the corresponding means are 49.07%, 50.00%, and 53.75%, with Ordinal Stage-State exceeding both alternatives in all three paired runs. The observed benefit differs across interface representations and training arrangements.

机器人操作阶段信息微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。