arXiv:2607.21227cs.RO2026-07

用冻结大模型指导力控装配,实现无损插接与自动恢复。

FORGE-plus: Force-Budgeted Recovery for Contact-Rich Assembly with a Frozen LLM Supervisor

论文配图:FORGE-plus: Force-Budgeted Recovery for Contact-Rich Assembly with a Frozen LLM Supervisor
图 1 · 摘自论文原文
  • 用文本大模型预设每件物品的力上限,通过力签名选恢复动作。
  • 256次测试零破损,平均峰值力仅5.4N,能准确判断释放时机。
  • 适合需要高精度力控的工业装配场景,尤其对易碎件友好。

力条件强化学习可在限定力矩下完成紧密配合装配,但实际应用需为每个物体设定合适力限,并在插入失败时恢复而不超限。本文提出两层框架:一个冻结的纯文本大语言模型(LLM)在执行前为每个物体分配力上限,并基于紧凑的力签名从固定动作库中选择恢复策略;LLM不直接控制力,由底层控制器强制执行力限,恢复策略无法提升力限,且断裂阈值仅评估者知晓。在两个夹爪(Robotiq 2F-140 和 Franka Panda 手)上测试脆弱瓶放置和0.4mm直径间隙齿轮插入任务。单一策略在256/256次评估中无破损,准确预测释放时机,并完成全流程拾取-插入,平均峰值力为5.4N。引入夹持滑移后,力签名恢复策略分别解决40%和64%的失败,而“更用力”基线要么无效,要么频繁导致破损。还报告了负面结果,包括PPO在严格力约束下无法完成任务,以及学习型释放策略失败。所有实验均在刚体仿真中进行,含隐藏断裂阈值;未做仿真到现实迁移声明。

原文摘要 · Abstract (English)

Force-conditioned reinforcement learning (RL) enables tight-clearance assembly under a commanded force ceiling, but practical deployment requires determining an appropriate force limit for each object and recovering from insertion failures without exceeding it. We present a two-layer framework in which a frozen, text-only large language model (LLM) assigns a per-object force ceiling before execution and selects recovery maneuvers from a fixed action menu using compact textual force signatures. The LLM never controls force directly: a low-level controller enforces the force ceiling, the recovery policy cannot increase it, and the hidden breaking-force threshold is known only to the evaluator. We evaluate the framework on fragile bottle placement and 0.4 mm diametral-clearance gear insertion using two grippers (Robotiq 2F-140 and Franka Panda hand). A single policy passes 256/256 evaluation episodes on both fragile and robust objects without breakage, correctly predicts release timing, and completes a full table-pick-and-insert pipeline with a mean peak force of 5.4 N. Under injected in-grip slip, the force-signature recovery strategy resolves 40% and 64% of failures on the two grippers, whereas a press-harder baseline is either ineffective or causes frequent breakage. We also report negative results, including the failure of PPO to solve the task under strict force constraints and unsuccessful learned release strategies. All experiments are conducted in rigid-body simulation with hidden force-threshold breakage; no sim-to-real claim is made.

力控装配大模型强化学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。