arXiv:2503.10509cs.LG2025-03被引 1

用大模型将强化学习行为转化为自然语言摘要,让智能体策略更易懂。

From Actions to Words: Towards Abstractive-Textual Policy Summarization in RL

  • 将智能体轨迹转为结构化文本,用大模型生成策略摘要。
  • 用户研究显示75.5%人更偏好其生成的摘要。
  • 无需任务微调,适用于长时序复杂环境,适合可解释性研究者。

解释强化学习智能体极具挑战,因其策略源于复杂的奖励结构与神经表示,人类难以理解。现有方法依赖人工演示,仅揭示局部行为,难以反映全局策略,用户需从原始观测推断意图。本文提出SySLLM(基于大语言模型的合成摘要),将策略解释重构为语言生成问题。不同于可视化演示,SySLLM将时空轨迹转化为结构化文本,通过提示大模型生成涵盖目标、探索风格与决策模式的连贯摘要。该框架无需任务特定微调,可扩展至长时程、语义丰富的环境,利用大模型的世界知识与组合推理能力,捕捉不同策略间的潜在行为结构。专家评估显示与人类分析高度一致,大规模用户研究发现75.5%参与者更倾向使用SySLLM摘要。结果表明,抽象文本摘要可成为解释复杂强化学习行为的新范式。

原文摘要 · Abstract (English)

Explaining reinforcement learning agents is challenging because policies emerge from complex reward structures and neural representations that are difficult for humans to interpret. Existing approaches often rely on curated demonstrations that expose local behaviors but provide limited insight into an agent's global strategy, leaving users to infer intent from raw observations. We propose SySLLM (Synthesized Summary using Large Language Models), a framework that reframes policy interpretation as a language-generation problem. Instead of visual demonstrations, SySLLM converts spatiotemporal trajectories into structured text and prompts an LLM to generate coherent summaries describing the agent's goals, exploration style, and decision patterns. SySLLM scales to long-horizon, semantically rich environments without task-specific fine-tuning, leveraging LLM world knowledge and compositional reasoning to capture latent behavioral structure across policies. Expert evaluations show strong alignment with human analyses, and a large-scale user study found that 75.5% of participants preferred SySLLM summaries over state-of-the-art demonstration-based explanations. Together, these results position abstractive textual summarization as a paradigm for interpreting complex RL behavior.

RL解释大模型文本生成策略摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。