arXiv:2504.17704cs.CL2025-04EMNLP综述被引 75

系统梳理大模型推理中的安全风险与防护方法

Safety in Large Reasoning Models: A Survey

  • 构建安全风险、攻击与防御的分类体系
  • 揭示推理模型在数学编码任务中的潜在漏洞
  • 适合关注AI安全与可信部署的研究者阅读

大型推理模型(LRMs)在数学和编程等任务中展现出卓越能力,依赖其先进的推理性能。然而,随着能力提升,其安全漏洞日益突出,可能影响实际应用部署。本文对LRMs进行系统性综述,深入分析新出现的安全风险、攻击手段及防御策略。通过建立详尽的分类框架,旨在清晰呈现当前LRMs的安全态势,为未来研究与开发提供结构化参考,以增强这些强大模型的安全性与可靠性。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) have exhibited extraordinary prowess in tasks like mathematics and coding, leveraging their advanced reasoning capabilities. Nevertheless, as these capabilities progress, significant concerns regarding their vulnerabilities and safety have arisen, which can pose challenges to their deployment and application in real-world settings. This paper presents a comprehensive survey of LRMs, meticulously exploring and summarizing the newly emerged safety risks, attacks, and defense strategies. By organizing these elements into a detailed taxonomy, this work aims to offer a clear and structured understanding of the current safety landscape of LRMs, facilitating future research and development to enhance the security and reliability of these powerful models.

大模型安全推理模型风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。