用自编码器分析大模型多智能体强化学习行为,发现隐藏策略与漏洞。
Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning
- 通过预训练自编码器和语言模型,提取可解释的行为特征。
- 90%的特征显著,发现角色扮演、语言切换及意外奖励作弊行为。
- 部分特征能提升下游任务表现,适合研究可信大模型行为的学者。
大型语言模型(LLMs)在复杂多智能体强化学习环境中应用日益广泛,但其行为演化难以理解。稀疏自编码器(SAEs)近期展现出数据驱动可解释性潜力。本文针对高复杂度环境Full-Press Diplomacy中的大规模强化学习训练过程,应用预训练SAEs并结合LLM摘要方法,提出Meta-Autointerp:将SAE特征聚类为关于训练动态的可解释假设。我们发现细粒度行为包括角色扮演模式、退化输出、语言切换,以及高层战略行为与环境特有漏洞。自动化评估验证90%的SAE元特征具有显著性,并发现令人意外的奖励作弊行为。然而,两个用户研究表明,即使看似有趣且有帮助的SAE特征,对人类而言也可能毫无价值甚至有害,大多数LLM生成的假设亦然。不过,一小部分由SAE推导出的假设在下游任务中具备预测能力。我们进一步通过增强未训练智能体的系统提示,使得分提升+14.2%。总体而言,SAEs与LLM摘要提供互补视角,共同构成未来数据驱动可解释性研究的实用起点,以确保大模型在整个训练过程中行为可信。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly trained in complex Reinforcement Learning, multi-agent environments, making it difficult to understand how behavior changes over training. Sparse Autoencoders (SAEs) have recently shown to be useful for data-centric interpretability. In this work, we analyze large-scale reinforcement learning training runs from the sophisticated environment of Full-Press Diplomacy by applying pretrained SAEs, alongside LLM-summarizer methods. We introduce Meta-Autointerp, a method for grouping SAE features into interpretable hypotheses about training dynamics. We discover fine-grained behaviors including role-playing patterns, degenerate outputs, language switching, alongside high-level strategic behaviors and environment-specific bugs. Through automated evaluation, we validate that 90% of discovered SAE Meta-Features are significant, and find a surprising reward hacking behavior. However, through two user studies, we find that even subjectively interesting and seemingly helpful SAE features may be worse than useless to humans, along with most LLM generated hypotheses. However, a subset of SAE-derived hypotheses are predictively useful for downstream tasks. We further provide validation by augmenting an untrained agent's system prompt, improving the score by +14.2%. Overall, we show that SAEs and LLM-summarizer provide complementary views into agent behavior, and together our framework forms a practical starting point for future data-centric interpretability work on ensuring trustworthy LLM behavior throughout training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。