arXiv:2505.10670cs.AIcs.CY2025-05被引 4

通过可解释特征调控,显著降低大模型代理的背叛概率。

Interpretable Risk Mitigation in LLM Agent Systems

  • 用稀疏自编码器提取可解释特征,操控残差流引导行为
  • 使背叛概率平均降低28个百分点,效果稳定可复现
  • 适合关注AI安全与可控性的研究者和开发者

由大语言模型驱动的自主代理在需负责任行动的领域展现出新应用前景,但其内在不可预测性带来了可靠性担忧。本文在基于改进版重复囚徒困境的游戏环境里,探索代理行为。提出一种与游戏机制和提示无关的策略修改方法:通过稀疏自编码器潜空间提取的可解释特征,调控残差流。使用善意协商特征进行引导,使平均背叛概率降低28个百分点。同时确定了多个开源大模型代理的可行调控范围。最后提出假设:将博弈论评估与表征调控对齐,有望推广至终端设备及具身平台的真实场景应用。

原文摘要 · Abstract (English)

Autonomous agents powered by large language models (LLMs) enable novel use cases in domains where responsible action is increasingly important. Yet the inherent unpredictability of LLMs raises safety concerns about agent reliability. In this work, we explore agent behaviour in a toy, game-theoretic environment based on a variation of the Iterated Prisoner's Dilemma. We introduce a strategy-modification method-independent of both the game and the prompt-by steering the residual stream with interpretable features extracted from a sparse autoencoder latent space. Steering with the good-faith negotiation feature lowers the average defection probability by 28 percentage points. We also identify feasible steering ranges for several open-source LLM agents. Finally, we hypothesise that game-theoretic evaluation of LLM agents, combined with representation-steering alignment, can generalise to real-world applications on end-user devices and embodied platforms.

大模型安全可解释性行为控制博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。