arXiv:2603.00719cs.RO2026-03被引 1

用关键帧引导奖励,提升实验室机器人长程操作成功率

Keyframe-Guided Structured Rewards for Reinforcement Learning in Long-Horizon Laboratory Robotics

  • 从示范中自动提取关键动作帧,生成分阶段目标
  • 40-60分钟内达成82%成功率,优于基线方法
  • 适合需要精确步骤逻辑的自动化实验机器人

实验室自动化中的长时序高精度操作(如移液枪头安装与液体转移)需在连续高维状态空间中遵守严格流程逻辑。现有方法面临奖励稀疏、多阶段结构约束及示范噪声问题,导致探索效率低、收敛不稳定。本文提出一种关键帧引导奖励生成框架:自动从示范中提取运动学感知的关键帧,通过潜在空间扩散预测器生成阶段目标,并构建基于几何进度的奖励机制以指导在线强化学习。该框架结合多视角视觉编码、潜在相似性进度追踪,以及基于视觉-语言-动作骨干网络的人机协同强化微调,使策略优化对齐生物实验流程的内在分步逻辑。在四个真实实验室任务中,包括高精度移液枪安装与动态液体转移,本方法经40–60分钟在线微调后平均成功率达82%。相比HG-DAgger(42%)和Hil-ConRFT(47%),证明了结构化关键帧引导奖励在克服探索瓶颈方面的有效性,为高精度长时序机器人实验室自动化提供可扩展解决方案。

原文摘要 · Abstract (English)

Long-horizon precision manipulation in laboratory automation, such as pipette tip attachment and liquid transfer, requires policies that respect strict procedural logic while operating in continuous, high-dimensional state spaces. However, existing approaches struggle with reward sparsity, multi-stage structural constraints, and noisy or imperfect demonstrations, leading to inefficient exploration and unstable convergence. We propose a Keyframe-Guided Reward Generation Framework that automatically extracts kinematics-aware keyframes from demonstrations, generates stage-wise targets via a diffusion-based predictor in latent space, and constructs a geometric progress-based reward to guide online reinforcement learning. The framework integrates multi-view visual encoding, latent similarity-based progress tracking, and human-in-the-loop reinforcement fine-tuning on a Vision-Language-Action backbone to align policy optimization with the intrinsic stepwise logic of biological protocols. Across four real-world laboratory tasks, including high-precision pipette attachment and dynamic liquid transfer, our method achieves an average success rate of 82% after 40--60 minutes of online fine-tuning. Compared with HG-DAgger (42%) and Hil-ConRFT (47%), our approach demonstrates the effectiveness of structured keyframe-guided rewards in overcoming exploration bottlenecks and providing a scalable solution for high-precision, long-horizon robotic laboratory automation.

机器人控制强化学习实验室自动化关键帧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。