arXiv:2604.27699cs.AI2026-04

让智能体基于价值观主动规划长期行为,突破被动执行局限。

Bridging Values and Behavior: A Hierarchical Framework for Proactive Embodied Agents

论文配图:Bridging Values and Behavior: A Hierarchical Framework for Proactive Embodied Agents
图 1 · 摘自论文原文
  • 用大模型推理抽象价值权衡,生成符号化子目标
  • 通过PDDL规划器将目标转为可执行动作,闭环反馈优化
  • 评估体系关注价值积累与行为多样性,适合自主系统研究者

当前具身智能体多局限于被动响应指令或即时需求满足,缺乏支撑长期自驱行为与解决动机冲突的稳定高层价值框架。我们提出《ValuePlanner》——一种分层认知架构,将高层价值调度与底层动作执行解耦。该架构采用基于大语言模型的认知模块,通过抽象价值权衡推理生成符号化子目标,并由经典PDDL规划器将其转化为可执行动作计划,过程通过闭环反馈机制优化。为衡量此类自主性,我们设计以价值为核心的评估套件,包含累积价值增益、偏好对齐度与行为多样性指标。在TongSim家庭环境中的实验表明,ValuePlanner能有效调和竞争性价值,生成连贯、长时程、自驱动的行为,优于指令跟随与需求驱动基线。本工作为连接内在价值观与具体行为提供了结构化路径。

原文摘要 · Abstract (English)

Current embodied agents are often limited to passive instruction-following or reactive need-satisfaction, lacking a stable, high-order value framework essential for long-term, self-directed behavior and resolving motivational conflicts. We introduce \textit{ValuePlanner}, a hierarchical cognitive architecture that decouples high-level value scheduling from low-level action execution. \textit{ValuePlanner} employs an LLM-based cognitive module to generate symbolic subgoals by reasoning through abstract value trade-offs, which are then translated into executable action plans by a classical PDDL planner. This process is refined via a closed-loop feedback mechanism. Evaluating such autonomy requires methods beyond task-success rates, and we therefore propose a value-centric evaluation suite measuring cumulative value gain, preference alignment, and behavioral diversity. Experiments in the TongSim household environment demonstrate that \textit{ValuePlanner} arbitrates competing values to generate coherent, long-horizon, self-directed behavior absent from instruction-following and needs-driven baselines. Our work offers a structured approach to bridging intrinsic values and grounded behavior for autonomous agents.

具身智能体价值对齐自主规划大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。