arXiv:2506.05451cs.SEcs.AI2025-06EMNLP综述被引 11

首份聚焦大模型安全与可解释性结合的综述,梳理70项关键研究。

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

  • 按大模型工作流程阶段构建可解释性方法新分类体系
  • 揭示70篇相关研究在安全改进中的具体作用路径
  • 适合关注大模型安全落地的研究者与开发者

随着大语言模型在现实世界中应用日益广泛,理解并缓解其不安全行为至关重要。可解释性技术能揭示不安全输出的原因并指导安全改进,但以往综述常忽视二者关联。本文首次系统弥合这一空白,提出一个统一框架,连接以安全为导向的可解释方法、其所支持的安全增强措施及其实际工具实现。基于大模型工作流程阶段的新分类体系,总结了近70项相关研究的关键交叉点。最后讨论开放挑战与未来方向。本综述及时帮助研究者与实践者掌握提升大模型安全性与可解释性的核心进展。

原文摘要 · Abstract (English)

As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical. Interpretation techniques can reveal causes of unsafe outputs and guide safety, but such connections with safety are often overlooked in prior surveys. We present the first survey that bridges this gap, introducing a unified framework that connects safety-focused interpretation methods, the safety enhancements they inform, and the tools that operationalize them. Our novel taxonomy, organized by LLM workflow stages, summarizes nearly 70 works at their intersections. We conclude with open challenges and future directions. This timely survey helps researchers and practitioners navigate key advancements for safer, more interpretable LLMs.

大模型安全可解释性综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。