让机器人学习更高效:主动挑选关键经验。
E$^2$DT: Efficient and Effective Decision Transformer with Experience-Aware Sampling for Robotic Manipulation

- 用决策变换器引导的智能采样,自动选最有价值数据。
- 在仿真和真实机器人上均超越已有方法,提升任务成功率。
- 适合做长序列机器人操控的科研与工程人员参考。
在机器人操作的强化学习中,决策变换器(DT)已成为应对长时序任务的有效框架。然而,其性能高度依赖于经验数据的覆盖范围。缺乏主动探索机制时,标准DT依赖均匀回放,导致样本效率低、探索受限、整体效果差。过度探索虽可避免局部最优,却常延迟策略收敛并降低效率。为此,我们提出E$^2$DT,一种基于体验感知采样的DT引导框架。该框架利用k-确定性点过程,使模型能主动塑造经验选择。它兼具高效性(优先高质量轨迹,如高回报、高不确定性、稀疏轨迹)与有效性(通过轨迹窗口间多样性保持策略最优性)。具体而言,DT内部隐状态衡量轨迹窗口间的多样性,质量则通过融合回报预期(RTG)分位数、预测不确定性及逆频率阶段覆盖率的复合指标量化。二者结合形成新颖的质量-多样性联合核函数,优先选取最具信息量的经验,实现高效且有效的学习。我们在仿真与真实机器人上的复杂操作基准测试中评估E$^2$DT,结果表明其持续优于先前方法。这些发现表明,将策略学习与体验感知采样相结合,为鲁棒的长时序机器人学习提供了一条系统化路径。
原文摘要 · Abstract (English)
In reinforcement learning (RL) for robotic manipulation, the Decision Transformer (DT) has emerged as an effective framework for addressing long-horizon tasks. However, DT's performance depends heavily on the coverage of collected experiences. Without an active exploration mechanism, standard DT relies on uniform replay, which leads to poor sample efficiency, limited exploration, and reduced overall effectiveness. At the same time, while excessive exploration can help avoid local optima, it often delays policy convergence and leads to degraded efficiency. To address these limitations, we propose E$^2$DT, a DT-guided k-Determinantal Point Process sampling framework that enables the model to actively shape its own experience selection. Our framework is experience-aware, allowing E$^2$DT to be both efficient, by prioritizing sampling quality, such as high-return, high-uncertainty, and underrepresented trajectories, and effective, by ensuring diversity across trajectory windows to preserve policy optimality. Specifically, DT's internal latent embeddings measure diversity across trajectory windows, while quality is quantified through a composite metric that integrates return-to-go (RTG) quantiles, predictive uncertainty, and stage coverage based on inverse frequency. These two dimensions are integrated into a novel quality-diversity joint kernel that prioritizes the most informative experiences, thereby enabling learning that is both efficient and effective. We evaluate E$^2$DT on challenging robotic manipulation benchmarks in both simulation and real-robot settings. Results show that it consistently outperforms prior methods. These findings demonstrate that coupling policy learning with experience-aware sampling provides a principled path toward robust long-horizon robotic learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。