让人类在推理时引导生成策略,不改模型也能精准控制行为轨迹。
Inference-Time Policy Steering through Human Interactions
- 用人类交互信号调整采样过程,不微调模型本身。
- 扩散采样策略在对齐意图与避免偏差间表现最佳。
- 适合需要实时人工干预的机器人长程任务场景。
通过人类示范训练的生成式策略可自主完成多模式、长时序任务。然而在推理阶段,人类通常被排除在执行循环之外,难以引导预训练策略选择特定子目标或轨迹形状。直接的人类干预可能加剧分布偏移,导致约束违反或执行失败。为此,我们提出推理时策略引导(ITPS)框架,利用人类交互来偏置生成采样过程,而非在交互数据上微调策略。我们在三个模拟和真实世界基准上评估了三种人类交互形式及对应的对齐距离度量。在六种采样策略中,所提出的带扩散的随机采样在对齐度与分布偏移之间取得最佳平衡。视频演示见 https://yanweiw.github.io/itps/。
原文摘要 · Abstract (English)
Generative policies trained with human demonstrations can autonomously accomplish multimodal, long-horizon tasks. However, during inference, humans are often removed from the policy execution loop, limiting the ability to guide a pre-trained policy towards a specific sub-goal or trajectory shape among multiple predictions. Naive human intervention may inadvertently exacerbate distribution shift, leading to constraint violations or execution failures. To better align policy output with human intent without inducing out-of-distribution errors, we propose an Inference-Time Policy Steering (ITPS) framework that leverages human interactions to bias the generative sampling process, rather than fine-tuning the policy on interaction data. We evaluate ITPS across three simulated and real-world benchmarks, testing three forms of human interaction and associated alignment distance metrics. Among six sampling strategies, our proposed stochastic sampling with diffusion policy achieves the best trade-off between alignment and distribution shift. Videos are available at https://yanweiw.github.io/itps/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。