arXiv:2605.08571cs.RO2026-05

用少量目标数据提升机器人生成策略的跨域适应能力

BEACON: Cross-Domain Co-Training of Generative Robot Policies via Best-Effort Adaptation

论文配图:BEACON: Cross-Domain Co-Training of Generative Robot Policies via Best-Effort Adaptation
图 1 · 摘自论文原文
  • 基于重要性重加权思想,动态调整源域数据贡献度
  • 在模拟到真实场景中实现更高鲁棒性与数据效率
  • 无需显式对齐也能自动实现特征对齐,适合多源学习

我们提出BEACON——一种基于最优适应的跨域协同训练框架,用于在大量源域演示和少量目标域演示下训练生成式机器人策略。该方法将跨域协同训练建模为一种偏差感知的重要性重加权问题,联合学习基于扩散模型的视觉-运动策略与每样本的源权重,以最小化受目标域泛化保证启发的目标函数。为实现高维序列策略的最佳努力适应,我们开发了可扩展的实例级偏差估计器、策略与权重的随机交替更新机制,以及支持异构源域的多源扩展方案。在模拟到模拟、模拟到真实及多源操控设置中,BEACON在鲁棒性和数据效率上均优于仅使用目标数据、固定比例协同训练及特征对齐基线。重要的是,即使没有显式对齐目标,BEACON仍能通过偏差感知的跨域协同训练隐式实现特征对齐。

原文摘要 · Abstract (English)

We introduce BEACON--Best-Effort Adaptation for Cross-Domain Co-Training--a theory-driven framework for training generative robot policies with abundant source demonstrations and limited target demonstrations. BEACON casts cross-domain co-training as a discrepancy-aware importance-reweighting problem, jointly learning a diffusion-based visuomotor policy and per-sample source weights that minimize an objective informed by target-domain generalization guarantees. To make best-effort adaptation practical for high-dimensional sequence policies, we develop scalable instance-level discrepancy estimators, stochastic alternating updates for policy and weights, and a multi-source extension that balances heterogeneous source domains. Across sim-to-sim, sim-to-real, and multi-source manipulation settings, BEACON improves robustness and data efficiency over target-only, fixed-ratio co-training, and feature-alignment baselines. Importantly, even without an explicit alignment objective, BEACON achieves feature alignment as an implicit result of discrepancy-aware cross-domain co-training.

机器人策略跨域学习生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。