arXiv:2505.18658cs.CLcs.AI2025-05中稿 · TMLR综述被引 16

系统梳理大模型鲁棒性问题与应对策略,助力AI更可靠。

Robustness in Large Language Models: A Survey of Mitigation Strategies and Evaluation Metrics

  • 从模型、数据、攻击三方面分析大模型不鲁棒的根源
  • 总结当前主流防御方法与评估指标体系
  • 适合关注AI可靠性与安全的研究者和开发者

大语言模型(LLMs)已成为自然语言处理与人工智能发展的关键支柱。然而,确保其鲁棒性仍是重大挑战。本文综述了该领域的最新研究进展:首先系统分析了大模型鲁棒性的概念基础,包括对多样化输入下性能一致性的重要性及实际应用中失效模式的影响;其次,从模型内在局限、数据驱动漏洞和外部对抗因素三方面归纳非鲁棒性来源;接着,回顾了当前主流的缓解策略,并讨论了广泛应用的基准测试、新兴评估指标以及评估真实可靠性方面的持续空白;最后,整合现有综述与跨学科研究,提炼趋势、未解问题与未来研究路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as a promising cornerstone for the development of natural language processing (NLP) and artificial intelligence (AI). However, ensuring the robustness of LLMs remains a critical challenge. To address these challenges and advance the field, this survey provides a comprehensive overview of current studies in this area. First, we systematically examine the nature of robustness in LLMs, including its conceptual foundations, the importance of consistent performance across diverse inputs, and the implications of failure modes in real-world applications. Next, we analyze the sources of non-robustness, categorizing intrinsic model limitations, data-driven vulnerabilities, and external adversarial factors that compromise reliability. Following this, we review state-of-the-art mitigation strategies, and then we discuss widely adopted benchmarks, emerging metrics, and persistent gaps in assessing real-world reliability. Finally, we synthesize findings from existing surveys and interdisciplinary studies to highlight trends, unresolved issues, and pathways for future research.

大模型鲁棒性评估指标安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。