arXiv:2511.22555cs.RO2025-11

让机器人不仅会完成任务,还做得优雅,通过实时干预优化动作质量。

Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention

  • 分离式优化框架,不重训基础策略,仅提升执行质量。
  • 引入隐式任务约束,用离线校准强化学习训练优雅度评判器。
  • 推理时仅在关键节点干预,适合追求高精度操作的场景。

视觉-语言-动作(VLA)模型推动了通用机器人操作的发展,但其策略执行质量参差不齐。我们归因于人类示范数据的质量混合性,即动作执行中的隐式原则仅部分满足。为此,我们提出LIBERO-Elegant基准,明确评估执行质量的标准。基于此,构建解耦式优化框架,在不修改或重新训练基础VLA策略的前提下,提升执行质量。我们将优雅执行形式化为对隐式任务约束(ITCs)的满足,并通过离线校准Q学习训练一个优雅度评判器,以估计候选动作的预期质量。推理时,即时干预(JITI)机制根据评判器置信度,在决策关键点进行选择性、按需修正。在LIBERO-Elegant及真实世界操作任务上的实验表明,该方法显著提升执行质量,甚至在未见过的任务上也有效。所提模型使机器人控制不仅关注任务成败,更重视执行过程的优雅性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have enabled notable progress in general-purpose robotic manipulation, yet their learned policies often exhibit variable execution quality. We attribute this variability to the mixed-quality nature of human demonstrations, where the implicit principles that govern how actions should be carried out are only partially satisfied. To address this challenge, we introduce the LIBERO-Elegant benchmark with explicit criteria for evaluating execution quality. Using these criteria, we develop a decoupled refinement framework that improves execution quality without modifying or retraining the base VLA policy. We formalize Elegant Execution as the satisfaction of Implicit Task Constraints (ITCs) and train an Elegance Critic via offline Calibrated Q-Learning to estimate the expected quality of candidate actions. At inference time, a Just-in-Time Intervention (JITI) mechanism monitors critic confidence and intervenes only at decision-critical moments, providing selective, on-demand refinement. Experiments on LIBERO-Elegant and real-world manipulation tasks show that the learned Elegance Critic substantially improves execution quality, even on unseen tasks. The proposed model enables robotic control that values not only whether tasks succeed, but also how they are performed.

机器人操作优雅执行实时干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。