arXiv:2506.11111cs.CLcs.AI2025-06综述被引 4

系统梳理大模型鲁棒性问题与评估方法,助力构建更稳定可靠的AI系统。

Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions

  • 按输入扰动类型分类,梳理对抗鲁棒性、分布外鲁棒性与评估方法
  • 总结最新评测数据集、指标与工具,覆盖噪声提示、幻觉等实际场景
  • 适合关注大模型安全与可靠性研究的开发者和研究人员参考

大语言模型(LLMs)因其强大的自然语言理解与生成能力近年来备受关注。随着其在智能体、具身智能等广泛场景中的应用,模型鲁棒性成为关键挑战。作为诸多AI应用的核心,LLMs需在面对恶意提示、低质量数据、分布外输入等意外情况时仍能保持输出一致性、正确性与稳定性。本文综述了大模型鲁棒性的核心概念与方法,首先给出鲁棒性的正式定义并说明综述收集标准;随后从输入扰动类型出发,分为三方面:1)对抗鲁棒性,应对故意操纵的提示(如噪声提示、长上下文、数据攻击等);2)分布外(OOD)鲁棒性,处理真实世界中未预期的应用场景(如分布外检测、零样本迁移、幻觉等);3)鲁棒性评估,归纳新出现的评测数据集、度量指标与工具。通过分析代表性工作,本文探讨该领域的未来机遇与研究方向,并整理相关文献,提供可搜索的项目主页(https://github.com/zhangkunzk/Awesome-LLM-Robustness-papers)以支持社区发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have gained enormous attention in recent years due to their capability of understanding and generating natural languages. With the rapid development and wild-range applications (e.g., Agents, Embodied Intelligence), the robustness of LLMs has received increased attention. As the core brain of many AI applications, the robustness of LLMs requires that models should not only generate consistent contents, but also ensure the correctness and stability of generated content when dealing with unexpeted application scenarios (e.g., toxic prompts, limited noise domain data, outof-distribution (OOD) applications, etc). In this survey paper, we conduct a thorough review of the robustness of LLMs, aiming to provide a comprehensive terminology of concepts and methods around this field and facilitate the community. Specifically, we first give a formal definition of LLM robustness and present the collection protocol of this survey paper. Then, based on the types of perturbated inputs, we organize this survey from the following perspectives: 1) Adversarial Robustness: tackling the problem that prompts are manipulated intentionally, such as noise prompts, long context, data attack, etc; 2) OOD Robustness: dealing with the unexpected real-world application scenarios, such as OOD detection, zero-shot transferring, hallucinations, etc; 3) Evaluation of Robustness: summarizing the new evaluation datasets, metrics, and tools for verifying the robustness of LLMs. After reviewing the representative work from each perspective, we discuss and highlight future opportunities and research directions in this field. Meanwhile, we also organize related works and provide an easy-to-search project (https://github.com/zhangkunzk/Awesome-LLM-Robustness-papers) to support the community.

大模型鲁棒性评估方法安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。