在无专家数据且数据稀缺的场景下,实现视觉域迁移的高效模仿学习。
State-Conditional Adversarial Learning: An Off-Policy Visual Domain Transfer Method for End-to-End Imitation Learning
- 基于状态条件的对抗学习框架,通过判别器估计条件KL散度来对齐潜在分布。
- 在真实自动驾驶环境中,仅用少量目标域数据即实现稳定迁移与高样本效率。
- 适合缺乏专家示范数据的现实场景,如机器人操控、自动驾驶等复杂任务。
我们研究在真实且具有挑战性的设置下,端到端模仿学习中的视觉域迁移问题,其中目标域数据严格为离线策略、无专家标注且稀少。我们首先进行理论分析,表明目标域模仿损失可被源域损失加上源与目标观测模型之间的状态条件隐变量KL散度所上界。基于此结果,我们提出状态条件对抗学习(SCAL),一种基于判别器估计条件KL项的离线策略对抗框架,以状态为条件对齐潜在分布。在基于BARC-CARLA模拟器构建的视觉差异较大的自动驾驶环境中进行实验,结果表明SCAL实现了稳健的域迁移和出色的样本效率。
原文摘要 · Abstract (English)
We study visual domain transfer for end-to-end imitation learning in a realistic and challenging setting where target-domain data are strictly off-policy, expert-free, and scarce. We first provide a theoretical analysis showing that the target-domain imitation loss can be upper bounded by the source-domain loss plus a state-conditional latent KL divergence between source and target observation models. Guided by this result, we propose State- Conditional Adversarial Learning, an off-policy adversarial framework that aligns latent distributions conditioned on system state using a discriminator-based estimator of the conditional KL term. Experiments on visually diverse autonomous driving environments built on the BARC-CARLA simulator demonstrate that SCAL achieves robust transfer and strong sample efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。