不训练模型,通过组合多个策略提升机器人控制性能。
Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition
- 用凸组合方式融合多个预训练策略的分布得分,测试时优化决策。
- 在多个基准上实现超越单个策略的性能,最高提升18%成功率。
- 适用于扩散与流匹配模型,支持异构策略即插即用。
基于扩散的机器人控制模型(如视觉-语言-动作、视觉-动作策略)已展现出强大能力,但其发展受限于大规模交互数据集的高昂获取成本。本文提出一种无需额外训练的增强策略性能新范式。令人意外的是,组合策略可超越任一父策略。贡献有三:首先,理论证明多个扩散模型的分布得分进行凸组合,能获得更优的单步目标函数,且通过格伦沃尔不等式证明该优势可沿生成轨迹传播,带来系统性提升;其次,提出通用策略组合(GPC)方法,通过测试时凸组合多个预训练策略的分布得分并结合搜索优化,实现性能增强;第三,通过Robomimic、PushT、RoboTwin等基准及真实机器人实验验证,GPC在多种任务中持续提升性能与适应性。对组合算子与权重策略的分析揭示了其成功机制。结果表明,GPC是一种简单有效的利用已有策略提升控制性能的方法。
原文摘要 · Abstract (English)
Diffusion-based models for robotic control, including vision-language-action (VLA) and vision-action (VA) policies, have demonstrated significant capabilities. Yet their advancement is constrained by the high cost of acquiring large-scale interaction datasets. This work introduces an alternative paradigm for enhancing policy performance without additional model training. Perhaps surprisingly, we demonstrate that the composed policies can exceed the performance of either parent policy. Our contribution is threefold. First, we establish a theoretical foundation showing that the convex composition of distributional scores from multiple diffusion models can yield a superior one-step functional objective compared to any individual score. A Grönwall-type bound is then used to show that this single-step improvement propagates through entire generation trajectories, leading to systemic performance gains. Second, motivated by these results, we propose General Policy Composition (GPC), a training-free method that enhances performance by combining the distributional scores of multiple pre-trained policies via a convex combination and test-time search. GPC is versatile, allowing for the plug-and-play composition of heterogeneous policies, including VA and VLA models, as well as those based on diffusion or flow-matching, irrespective of their input visual modalities. Third, we provide extensive empirical validation. Experiments on Robomimic, PushT, and RoboTwin benchmarks, alongside real-world robotic evaluations, confirm that GPC consistently improves performance and adaptability across a diverse set of tasks. Further analysis of alternative composition operators and weighting strategies offers insights into the mechanisms underlying the success of GPC. These results establish GPC as a simple yet effective method for improving control performance by leveraging existing policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。