arXiv:2607.25593cs.RO2026-07被引 1

旧数据何时有用?升级机器人后,能力达到阈值才开始见效。

When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning

论文配图:When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning
图 1 · 摘自论文原文
  • 发现跨配置数据在任务能力达阈值前无效,之后突然提升
  • 花插入任务中准确率从23.3%飙升至86.7%,笔插入任务增85.0%到93.3%
  • 提出三阶段模式,指导新旧数据收集时机

机器人硬件持续演进,但示范数据常绑定特定传感器与执行器配置。这引出一个未被充分探讨的实际问题:旧数据何时开始对升级后的机器人有帮助?我们在两代轮式人形平台间研究此问题,摄像头和夹爪均更换但整体形态保持不变。与普遍认为更多跨配置数据总有益的假设相反,我们观察到类似‘顿悟’的转变:旧数据在升级配置任务能力未达最低阈值前无用,一旦越过阈值,联合训练收益急剧上升,随后趋于饱和。我们提出该任务依赖性转变由转移阈值决定,并刻画出三阶段模式。在真实机器人操作任务中,三种阶段均出现:低能力时(10.0%→10.0%)无明显收益,越过阈值后(花插入:23.3%→86.7%),收益陡增;高能力时(笔插入:85.0%→93.3%)边际递减。基于梯度对齐与残余策略不确定性,我们提供理论解释,并推导出分阶段决策规则,用于判断何时应采集新硬件数据,何时可复用旧演示数据。该三阶段模式在移动双臂浇水任务中得到验证,结果与预测一致。

原文摘要 · Abstract (English)

Robotic hardware evolves over time, but demonstration data is often tied to a specific sensor and actuator configuration. This raises a practical and underexplored question: when does legacy data begin to benefit an upgraded robot? We study this question on a wheeled humanoid platform across two hardware generations, where both the camera and gripper are changed while the overall morphology remains fixed. Contrary to the common assumption that more cross-configuration data is always helpful, we observe a grokking-like transition: legacy data remains ineffective until the upgraded configuration acquires a minimum level of task competence, after which co-training gains rise sharply before diminishing near saturation. We hypothesize that this task-dependent transition is governed by a transfer threshold and characterize the resulting three-phase pattern. Across real-robot manipulation tasks, we observe all three phases: no measurable benefit at low competence ($10.0\% \rightarrow 10.0\%$), a sharp gain after crossing the threshold ($23.3\% \rightarrow 86.7\%$ on flower insertion), and diminishing returns at high competence ($85.0\% \rightarrow 93.3\%$ on pen insertion). We provide a theoretical account based on gradient alignment and residual policy uncertainty, and derive a phase-aware rule for deciding when to collect more new-hardware data and when to reuse legacy demonstrations. We further validate this three-phase pattern on a mobile dual-arm watering task, with results consistent with our predictions.

机器人学习迁移学习数据复用三阶段模式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。