arXiv:2603.22813cs.AI2026-03

让智能体动态推断偏好,适应目标变化

Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts

  • 用概率信念追踪隐藏的偏好权重并实时更新
  • 在目标突变后性能优于固定权重方法
  • 适合目标随环境变化的复杂决策场景

人类在不同情境下会调整优先级,而非遵循固定目标。现有强化学习方法多假设偏好权重恒定。本文提出动态偏好推断(DPI)框架,将偏好权重视为随上下文漂移的隐变量。智能体通过近期交互数据更新对偏好的概率信念,并据此调整策略。在队列、迷宫和连续控制任务中,使用向量回报作为隐性权衡证据,联合训练变分偏好推断模块与偏好条件策略网络。DPI 在目标突发变化后能快速适应,性能显著优于固定权重和启发式基线方法。

原文摘要 · Abstract (English)

Humans often juggle multiple, sometimes conflicting objectives and shift their priorities as circumstances change, rather than following a fixed objective function. In contrast, most computational decision-making and multi-objective RL methods assume static preference weights or a known scalar reward. In this work, we study sequential decision-making problem when these preference weights are unobserved latent variables that drift with context. Specifically, we propose Dynamic Preference Inference (DPI), a cognitively inspired framework in which an agent maintains a probabilistic belief over preference weights, updates this belief from recent interaction, and conditions its policy on inferred preferences. We instantiate DPI as a variational preference inference module trained jointly with a preference-conditioned actor-critic, using vector-valued returns as evidence about latent trade-offs. In queueing, maze, and multi-objective continuous-control environments with event-driven changes in objectives, DPI adapts its inferred preferences to new regimes and achieves higher post-shift performance than fixed-weight and heuristic envelope baselines.

强化学习动态偏好多目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。