arXiv:2608.27072cs.LGcs.AI2026-08

让智能体根据情绪自主调节目标优先级,提升复杂环境下的决策能力。

Emotional Preferences as Goal-Priority Regulation

论文配图:Emotional Preferences as Goal-Priority Regulation
图 1 · 摘自论文原文
  • 用强化学习构建内外双层框架,实现高阶目标驱动的情绪偏好自动生成。
  • 实验显示新方法可动态切换目标优先级,性能优于固定或人工设定偏好。
  • 适合研究智能体自主决策、情绪机制与多目标优化的学者参考。

智能体在决策中面临的核心问题是:竞争性低阶目标的相对优先级能否由高层目标自主生成的情绪偏好决定,而非外部预设。在外部环境变化和内部状态演化的背景下,情绪在调节各目标优先级方面起着关键功能作用。受目标导向情绪理论启发,本文探讨如何通过强化学习实现这种偏好调节的计算建模。我们提出一种涌现式情绪偏好概念:高层目标自主诱导对竞争性低阶目标的状态依赖偏好。该概念基于包含多目标强化学习内控制器与外偏好生成器的框架:内控制器提供一系列偏好条件下的目标导向行为库,外偏好生成器通过高层目标上的强化学习,学习从当前状态到目标偏好的映射。我们将情绪偏好操作化为通过优化过程产生的状态依赖型目标优先级调节。此外,我们刻画了偏好调节所诱导的策略空间,并推导出最优性差距的上界,其与内行为库的表征误差相关。当最优策略可由现有偏好条件策略表示时,该差距趋于零。在自建的多目标探索环境中实验表明,学习到的偏好函数展现出上下文相关的优先级切换、分级权衡及时间持续性,且优于评估的固定偏好和人工设计偏好策略。

原文摘要 · Abstract (English)

A core question in decision-making for agents is whether the relative priorities of competing lower-level objectives can be determined by emotional preferences autonomously generated by higher-level goals, rather than being externally prespecified. Under changing external environments and evolving internal states, emotions play an important functional role in regulating the relative priorities of competing goals. Inspired by the goal-directed theory of emotion, this paper studies how such preference regulation can be computationally realized through reinforcement learning. We first propose a conception of emergent emotional preference: a high-level goal autonomously induces state-dependent preferences over competing lower-level objectives. This conception is built upon a framework consisting of a multi-objective reinforcement learning inner controller and an outer preference generator. The inner controller provides a repertoire of preference-conditioned goal-directed behaviors, while the outer preference generator learns a mapping from the current state to objective preferences through reinforcement learning on a high-level goal. We operationalize emotional preference as a state-dependent regulation of relative goal priorities that emerges through optimization. Furthermore, we characterize the policy space induced by preference regulation and derive an upper bound on the optimality gap in terms of the representation error of the inner behavioral repertoire. We show that the gap vanishes when the optimal policy can be represented by the available preference-conditioned policies. Experiments in self-constructed multi-objective exploration environments show that the learned preference function exhibits contextual priority switching, graded trade-offs, and temporal persistence, and outperforms the evaluated fixed-preference and handcrafted-preference strategies.

强化学习情绪机制多目标决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。