arXiv:2502.10732cs.LGcs.AI2025-02被引 5

让语言模型同时生成可解释的决策规则并优化选择,提升资源分配的透明性与效率。

Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents

  • 用语言模型生成候选规则,通过强化学习筛选最优规则并生成推理链。
  • 在真实场景中性能接近深度强化学习,且解释质量更高。
  • 适合需要透明决策的医疗、公共政策等高风险领域应用。

深度强化学习在医疗、公共政策和资源管理等领域的序列资源分配问题中表现优异,但其决策过程缺乏透明性与适应性,难以与人类协同。相比之下,基于大语言模型的语言代理虽能提供可理解的推理过程,但在决策有效性上存在不足。为此,本文提出规则瓶颈强化学习(RBRL)框架,联合优化决策与解释。每一步中,RBRL利用语言模型生成候选规则,通过基于注意力机制的强化学习策略进行选择,并结合思维链推理确定环境动作及解释。强化学习的规则选择同时依据环境奖励和由语言模型评估的可解释性指标进行优化。在真实场景中的评估表明,RBRL在性能上媲美深度强化学习,在效率上优于语言模型微调;问卷调查进一步验证了其解释质量的显著提升。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (RL) is remarkably effective in addressing sequential resource allocation problems in domains such as healthcare, public policy, and resource management. However, deep RL policies often lack transparency and adaptability, challenging their deployment alongside human decision-makers. In contrast, Language Agents, powered by large language models (LLMs), provide human-understandable reasoning but may struggle with effective decision making. To bridge this gap, we propose Rule-Bottleneck Reinforcement Learning (RBRL), a novel framework that jointly optimizes decision and explanations. At each step, RBRL generates candidate rules with an LLM, selects among them using an attention-based RL policy, and determines the environment action with an explanation via chain-of-thought reasoning. The RL rule selection is optimized using the environment rewards and an explainability metric judged by the LLM. Evaluations in real-world scenarios highlight RBRL's competitive performance with deep RL and efficiency gains over LLM fine-tuning. A survey further confirms the enhanced quality of its explanations.

强化学习可解释性语言模型资源分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。