arXiv:2602.09580cs.ROcs.LG2026-02被引 2

用流模型提升真实机器人灵巧操作的样本效率与稳定性。

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows

  • 采用流模型生成多模态动作,实现精确似然更新。
  • 动作分块批评器改善长时序信用分配,提升适应性。
  • 适合需高精度控制的复杂真实机器人任务。

真实世界中灵巧操作策略的微调面临交互预算有限和动作分布高度多模态的挑战。基于扩散模型的策略虽表达能力强,但因动作概率不可计算,无法进行保守的似然更新。传统高斯策略在动作分块执行时易坍塌,且标准逐步批评器无法匹配分块执行结构,导致信用分配不佳。本文提出SERFN,一种基于归一化流(NF)的样本高效离策略微调框架。该框架通过归一化流策略获得多模态动作段的精确似然,支持保守、稳定的策略更新,提升样本效率;同时引入动作分块批评器,评估完整动作序列,使价值估计与策略的时间结构对齐,改善长时序信用分配。据我们所知,这是首个在真实机器人硬件上实现基于似然的多模态生成策略与分块级价值学习结合的实例。我们在两个具有挑战性的现实任务上评估:从盒中取剪刀剪胶带,以及掌心朝下抓持下旋转立方体——两者均需长时间、高精度的灵巧控制。实验表明,SERFN在这些任务上实现了稳定、高效的适应,而标准方法难以奏效。

原文摘要 · Abstract (English)

Real-world fine-tuning of dexterous manipulation policies remains challenging due to limited real-world interaction budgets and highly multimodal action distributions. Diffusion-based policies, while expressive, do not permit conservative likelihood-based updates during fine-tuning because action probabilities are intractable. In contrast, conventional Gaussian policies collapse under multimodality, particularly when actions are executed in chunks, and standard per-step critics fail to align with chunked execution, leading to poor credit assignment. We present SERFN, a sample-efficient off-policy fine-tuning framework with normalizing flow (NF) to address these challenges. The normalizing flow policy yields exact likelihoods for multimodal action chunks, allowing conservative, stable policy updates through likelihood regularization and thereby improving sample efficiency. An action-chunked critic evaluates entire action sequences, aligning value estimation with the policy's temporal structure and improving long-horizon credit assignment. To our knowledge, this is the first demonstration of a likelihood-based, multimodal generative policy combined with chunk-level value learning on real robotic hardware. We evaluate SERFN on two challenging dexterous manipulation tasks in the real world: cutting tape with scissors retrieved from a case, and in-hand cube rotation with a palm-down grasp -- both of which require precise, dexterous control over long horizons. On these tasks, SERFN achieves stable, sample-efficient adaptation where standard methods struggle.

灵巧操作强化学习归一化流机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。