arXiv:2511.03181cs.ROcs.LG2025-11

用AI统一控制机器人与人协作包纸,成功率97%。

Learning-based Cooperative Robotic Paper Wrapping: A Unified Control Policy with Residual Force Control

  • 用大模型规划任务,用混合学习策略控制动作
  • 97%成功率,能自动处理整套包纸流程
  • 适合需要人机协同的柔性物体操作场景

人机协作在仓库和零售环境中至关重要,尤其在处理纸张、袋子等易变形物体时。为应对这一挑战,本文聚焦礼品包装任务,该任务涉及长时序操控、精确折叠、可控折痕和牢固固定,成功标准为包装整洁无破损。提出一种基于学习的框架,结合大语言模型(LLM)驱动的高层任务规划与低层混合模仿学习(IL)与强化学习(RL)策略。核心是子任务感知的机器人变换器(START),通过人类示范学习统一策略,显式引入子任务标识实现时间定位,捕捉全序列的长期依赖关系。相比传统方法,该策略不依赖短任务分块,而是学习子目标,具备更强鲁棒性与灵活性。实测中,该框架在真实包装任务上达到97%成功率,减少对专用模型的需求,支持可控的人类监督,并有效连接高层意图与细粒度力控需求。

原文摘要 · Abstract (English)

Human-robot cooperation is essential in environments such as warehouses and retail stores, where workers frequently handle deformable objects like paper, bags, and fabrics. Coordinating robotic actions with human assistance remains difficult due to the unpredictable dynamics of deformable materials and the need for adaptive force control. To explore this challenge, we focus on the task of gift wrapping, which exemplifies a long-horizon manipulation problem involving precise folding, controlled creasing, and secure fixation of paper. Success is achieved when the robot completes the sequence to produce a neatly wrapped package with clean folds and no tears. We propose a learning-based framework that integrates a high-level task planner powered by a large language model (LLM) with a low-level hybrid imitation learning (IL) and reinforcement learning (RL) policy. At its core is a Sub-task Aware Robotic Transformer (START) that learns a unified policy from human demonstrations. The key novelty lies in capturing long-range temporal dependencies across the full wrapping sequence within a single model. Unlike vanilla Action Chunking with Transformer (ACT), typically applied to short tasks, our method introduces sub-task IDs that provide explicit temporal grounding. This enables robust performance across the entire wrapping process and supports flexible execution, as the policy learns sub-goals rather than merely replicating motion sequences. Our framework achieves a 97% success rate on real-world wrapping tasks. We show that the unified transformer-based policy reduces the need for specialized models, allows controlled human supervision, and effectively bridges high-level intent with the fine-grained force control required for deformable object manipulation.

人机协作柔性物体强化学习包装机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。