arXiv:2604.13645cs.ROcs.AI2026-04被引 5

解析仿真与真实数据联合训练的机制,揭示性能提升关键因素。

A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies

论文配图:A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies
图 1 · 摘自论文原文
  • 发现结构化表征对齐是核心机制,平衡跨域对齐与领域区分性。
  • 识别重要性重加权效应,由不同领域动作权重调节产生。
  • 实验验证机制有效性,适合研究机器人策略训练的学者参考。

联合训练结合有限的真实世界数据与大量替代数据(如仿真或跨实体机器人数据),被广泛用于生成式机器人策略的训练。尽管其在实践中表现良好,但决定何时及为何有效的内在机制仍不清晰。本文通过理论分析与实证研究,揭示了影响性能的两个内在效应:一是‘结构化表征对齐’,反映跨域表征对齐与领域可区分性的平衡,主导下游性能;二是‘重要性重加权效应’,源于领域相关的动作权重调制,在次要层面起作用。我们在简化模型上进行控制实验,并在广泛的模拟-模拟和模拟-真实机器人操作任务中验证这些效应。分析结果为近期联合训练技术提供了统一解释,并提出一种简单方法,能持续优于已有方案。本研究旨在剖析联合训练的内在机理,推动该方向的进一步探索。

原文摘要 · Abstract (English)

Co-training, which combines limited in-domain real-world data with abundant surrogate data such as simulation or cross-embodiment robot data, is widely used for training generative robot policies. Despite its empirical success, the mechanisms that determine when and why co-training is effective remain poorly understood. We investigate the mechanism of sim-and-real co-training through theoretical analysis and empirical study, and identify two intrinsic effects governing performance. The first, \textbf{``structured representation alignment"}, reflects a balance between cross-domain representation alignment and domain discernibility, and plays a primary role in downstream performance. The second, the \textbf{``importance reweighting effect"}, arises from domain-dependent modulation of action weighting and operates at a secondary level. We validate these effects with controlled experiments on a toy model and extensive sim-and-sim and sim-and-real robot manipulation experiments. Our analysis offers a unified interpretation of recent co-training techniques and motivates a simple method that consistently improves upon prior approaches. More broadly, our aim is to examine the inner workings of co-training and to facilitate research in this direction.

机器人联合训练生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。