arXiv:2605.10166cs.RO2026-05

用低质量轨迹训练3D机器人,通过想象和重排序提升成功率。

Data-Asymmetric Latent Imagination and Reranking for 3D Robotic Imitation Learning

论文配图:Data-Asymmetric Latent Imagination and Reranking for 3D Robotic Imitation Learning
图 1 · 摘自论文原文
  • 从混合质量轨迹中学习潜在世界模型,生成虚拟尝试路径。
  • 在不增加高质量数据前提下,平均提升6.8%任务成功率。
  • 适合数据稀缺或实验成本高的机器人学习场景。

机器人模仿学习通常依赖最优示范,但现实中采集的数据常包含次优、探索性甚至失败的轨迹。若直接丢弃这些数据,会浪费关于环境动态和失败模式的重要信息。虽然3D策略通过强空间泛化减少了对高质量示范的依赖,但仍需大规模数据才能实现高任务成功率。为此,我们提出DALI-R框架,利用3D点云构建潜在世界模型以进行虚拟推演,并设计任务完成评分器对候选动作片段进行重排序,从而在不引入额外高质量示范的前提下改进决策能力。该框架在扩散模型与高效流匹配策略上均被验证,于Adroit和MetaWorld基准测试中,两种3D基线策略均实现平均6.8%的成功率提升,推理开销增加不足0.7倍。

原文摘要 · Abstract (English)

Robotic imitation learning typically assumes access to optimal demonstrations, yet real-world data collection often yields suboptimal, exploratory, or even failed trajectories. Discarding such data wastes valuable information about environment dynamics and failure modes, which can instead be leveraged to improve decision-making. While 3D policies reduce reliance on high-quality demonstrations through strong spatial generalization, they still require large-scale data to achieve high task success. To address this, we propose DALI-R, a Data-Asymmetric Latent Imagination and Reranking framework for 3D robotic imitation learning from mixed-quality trajectories. It learns a Latent World Model over 3D point clouds for imagined rollouts and a Task Completion Scorer that reranks candidate action chunks, improving decision-making without additional high-quality demonstrations. We instantiate DALI-R with both diffusion and efficient flow-matching policies and evaluate it on Adroit and MetaWorld benchmarks. Across the two evaluated 3D base policies, DALI-R achieves an average $6.8$\% improvement in success rate while incurring less than $0.7\times$ additional inference overhead.

机器人学习3D模仿低质量数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。