用自然语言动态调整AI决策,兼顾任务效果与人类偏好。
VORTEX: Aligning Task Utility and Human Preferences through LLM-Guided Reward Shaping
- 通过大模型生成语言引导的奖励函数,动态优化决策
- 在真实分配任务中同时提升人类对齐度与任务性能
- 无需改写求解器或设定权重,适合需要人机协作的场景
在社会影响优化中,人工智能决策系统通常依赖可精准量化的数学目标进行求解。然而,这些系统难以直接融入以自然语言表达的动态人类偏好。近期方法尝试使用大语言模型(LLMs)从偏好描述生成新奖励函数,虽灵活但可能牺牲系统的根本性能保障。本文提出 exttt{VORTEX},一种语言引导的奖励塑造框架,在保持原有优化目标的同时,自适应融合人类反馈。我们将问题形式化为多目标优化,利用大模型基于言语强化与文本梯度提示迭代生成塑造奖励。这使利益相关方可通过自然语言引导决策行为,而无需修改求解器或指定权衡系数。我们提供了理论保证: exttt{VORTEX} 收敛于效用与偏好满足之间的帕累托最优折衷。实证结果表明,在真实分配任务中, exttt{VORTEX} 在满足人类对齐覆盖目标方面优于基线,同时保持高任务性能。本工作提出了一种兼具实用性与理论基础的人机协同优化范式。
原文摘要 · Abstract (English)
In social impact optimization, AI decision systems often rely on solvers that optimize well-calibrated mathematical objectives. However, these solvers cannot directly accommodate evolving human preferences, typically expressed in natural language rather than formal constraints. Recent approaches address this by using large language models (LLMs) to generate new reward functions from preference descriptions. While flexible, they risk sacrificing the system's core utility guarantees. In this paper, we propose \texttt{VORTEX}, a language-guided reward shaping framework that preserves established optimization goals while adaptively incorporating human feedback. By formalizing the problem as multi-objective optimization, we use LLMs to iteratively generate shaping rewards based on verbal reinforcement and text-gradient prompt updates. This allows stakeholders to steer decision behavior via natural language without modifying solvers or specifying trade-off weights. We provide theoretical guarantees that \texttt{VORTEX} converges to Pareto-optimal trade-offs between utility and preference satisfaction. Empirical results in real-world allocation tasks demonstrate that \texttt{VORTEX} outperforms baselines in satisfying human-aligned coverage goals while maintaining high task performance. This work introduces a practical and theoretically grounded paradigm for human-AI collaborative optimization guided by natural language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。