让机器人学会按用户偏好折叠布料,提升个性化操作能力。
Preference Aligned Visuomotor Diffusion Policies for Deformable Object Manipulation
- 结合RPO与KTO思想,用少量示范实现偏好对齐
- 在多种衣物折叠任务中表现更优且样本效率更高
- 适合需要个性化操作的柔性物体抓取场景
人类在执行操作任务时往往有隐含的偏好,这些偏好微妙、主观且难以言说。尽管机器人若能理解这些偏好可显著提升个性化程度与用户满意度,但在柔性物体(如衣物)操作中仍鲜有研究。本文提出RKO方法,通过少量示范将预训练的视觉-运动扩散策略适配至用户偏好。我们在真实世界中测试了多种衣物折叠任务,对比RKO与主流偏好学习框架(包括RPO、KTO)及基础扩散策略。结果表明,偏好对齐策略(尤其是RKO)在性能和样本效率上均优于标准微调方法。这验证了结构化偏好学习在复杂柔性物体操作中实现个性化行为的可行性和重要性。
原文摘要 · Abstract (English)
Humans naturally develop preferences for how manipulation tasks should be performed, which are often subtle, personal, and difficult to articulate. Although it is important for robots to account for these preferences to increase personalization and user satisfaction, they remain largely underexplored in robotic manipulation, particularly in the context of deformable objects like garments and fabrics. In this work, we study how to adapt pretrained visuomotor diffusion policies to reflect preferred behaviors using limited demonstrations. We introduce RKO, a novel preference-alignment method that combines the benefits of two recent frameworks: RPO and KTO. We evaluate RKO against common preference learning frameworks, including these two, as well as a baseline vanilla diffusion policy, on real-world cloth-folding tasks spanning multiple garments and preference settings. We show that preference-aligned policies (particularly RKO) achieve superior performance and sample efficiency compared to standard diffusion policy fine-tuning. These results highlight the importance and feasibility of structured preference learning for scaling personalized robot behavior in complex deformable object manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。