用生成式流模型构建奖励,让大模型更懂如何解释智能体决策。
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
- 用连续归一化流生成多元且概率化的解释奖励
- 解释准确率提升,逻辑性和可操作性更强
- 适合需要可靠解释的AI协作场景
随着人类与由强化学习、大语言模型等驱动的多样化智能体共存,以自然语言解释智能体策略的能力对可靠共存至关重要。我们提出一个通用框架,通过人工智能反馈强化学习训练生成解释的大模型,利用生成式连续归一化流(CNFs)生成分布式奖励。CNFs捕捉了人类对解释评价的多元性和概率特性。在温和假设下,当使用大模型生成的噪声代理奖励训练时,CNFs可证明地限制与真实人类奖励分布的偏差。我们设计了一种专用的CNF架构,能有选择地关注决策上下文和解释中的语言线索来生成奖励。人工及大模型评估均表明,相比代理大模型奖励或现有RLHF、RLAIF基线,本方法生成的解释能更准确预测真实智能体行为,具备更强逻辑性与可操作性,并降低认知负荷。
原文摘要 · Abstract (English)
As humans increasingly share environments with diverse agents powered by RL, LLMs, and beyond, the ability to explain agent policies in natural language is vital for reliable coexistence. We introduce a general-purpose framework that trains explanation-generating LLMs via reinforcement learning from AI feedback, with distributional rewards generated by generative continuous normalizing flows (CNFs). CNFs capture the pluralistic and probabilistic nature of human judgments about explanations. Moreover, under mild assumptions, CNFs provably bound deviations from true human reward distributions when trained on noisy proxy rewards from LLMs. We design a specialized CNF architecture that selectively attends to linguistic cues in the decision context and explanations when generating rewards. Human and LLM evaluators find that our method delivers explanations that enable more accurate predictions of true agent decisions, exhibit greater logical soundness and actionability, and impose lower cognitive load than explanations trained with proxy LLM rewards or state-of-the-art RLHF and RLAIF baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。