让机器人根据用户偏好动态调整规划保守性,无需重新训练。
MoMo: Conditioned Contrastive Representation Learning for Preference-Modulated Planning

- 用条件化对比学习构建可调节的潜在表示空间
- 同一任务下按用户偏好生成从激进到保守的多种路径
- 适合需要灵活安全策略的机器人规划场景
时序对比表示学习能将长程规划简化为低维线性系统的推理。但现有方法仅学习单一潜在几何结构,无法区分相同起终点下效率与风险权衡的不同有效行为。本文提出MoMo,一种偏好条件化的对比规划方法,通过标量用户偏好在推理时连续调节规划保守性,无需重训练。MoMo利用特征逐维线性调制和低秩神经调制,联合学习表示几何与潜在预测算子的条件化。其形式保持了表示空间中所需的概率密度比,确保推理效率。在六个环境中,MoMo能根据用户偏好平滑调节计划安全性,在时间一致性和偏好一致性上优于状态增强基线方法。
原文摘要 · Abstract (English)
Temporally contrastive representation learning induces a latent structure capable of reducing long-horizon planning to inference in a low-dimensional linear system. However, existing contrastive planning work learns a single latent geometry which cannot distinguish multiple valid behaviors trading task efficiency against risk exposure for the same start-goal query. We introduce MoMo, a preference-conditioned contrastive planner allowing a scalar user preference to continuously modulate plan conservativeness at inference time, without retraining. MoMo learns a joint conditioning of the representation geometry and latent prediction operator via Feature-Wise Linear Modulation and low-rank neural modulation, respectively. We show that our formulation preserves the probability density ratio encoded in the representation space that is required for inference-driven contrastive planning, further retaining its inference-time efficiency. Across six environments, MoMo smoothly adapts plan safety according to user preferences, yielding improved temporal and preferential consistency over state augmentation baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。