arXiv:2603.03280cs.ROcs.AI2026-03被引 1

用人类偏好优化刀工剥皮,让机器人学会像人一样判断操作好坏。

How to Peel with a Knife: Aligning Fine-Grained Manipulation with Human Preference

  • 分两阶段训练:先学基础刀法,再结合人类反馈调优质量。
  • 仅需50-200条数据,剥果蔬成功率超90%,提升40%。
  • 学会一种菜后能零样本应对新食材,适合复杂精细操作研究。

许多关键操作任务——如食品准备、手术和手工艺——仍难以由自主机器人完成。这些任务不仅具有接触密集、力觉敏感的动态特性,还存在‘隐含’的成功标准:与抓放任务不同,其任务质量是连续且主观的(例如土豆剥得是否干净),导致量化评估和奖励工程困难。本文以刀工剥皮为例,提出一种学习框架。方法分为两阶段:首先通过力觉感知的数据采集与模仿学习,获得鲁棒的初始策略,实现对物体变化的泛化;其次,利用结合定量指标与定性人类反馈的奖励模型,通过偏好式微调优化策略,使其行为更符合人类对任务质量的判断。仅使用50-200条剥皮轨迹,系统在黄瓜、苹果、土豆等挑战性食材上达到超过90%的平均成功率,偏好微调可使性能提升最高达40%。令人惊讶的是,仅在单一食材类别上训练的策略,对同类别未见实例及跨类别的分布外食材均表现出强零样本泛化能力,成功率仍保持在90%以上。

原文摘要 · Abstract (English)

Many essential manipulation tasks - such as food preparation, surgery, and craftsmanship - remain intractable for autonomous robots. These tasks are characterized not only by contact-rich, force-sensitive dynamics, but also by their "implicit" success criteria: unlike pick-and-place, task quality in these domains is continuous and subjective (e.g. how well a potato is peeled), making quantitative evaluation and reward engineering difficult. We present a learning framework for such tasks, using peeling with a knife as a representative example. Our approach follows a two-stage pipeline: first, we learn a robust initial policy via force-aware data collection and imitation learning, enabling generalization across object variations; second, we refine the policy through preference-based finetuning using a learned reward model that combines quantitative task metrics with qualitative human feedback, aligning policy behavior with human notions of task quality. Using only 50-200 peeling trajectories, our system achieves over 90% average success rates on challenging produce including cucumbers, apples, and potatoes, with performance improving by up to 40% through preference-based finetuning. Remarkably, policies trained on a single produce category exhibit strong zero-shot generalization to unseen in-category instances and to out-of-distribution produce from different categories while maintaining over 90% success rates.

机器人操作人类偏好精细动作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。