arXiv:2502.09673cs.CLcs.AI2025-02被引 21

提升大模型推理能力可能带来隐藏安全风险,但也可用于增强安全性。

Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning

  • 通过提示和微调增强模型推理能力
  • 发现推理越强,潜在安全风险越高
  • 利用推理机制反向提升模型安全性

大型语言模型(LLMs)在各类自然语言处理基准测试中表现出色。然而,要胜任需要精细推理与精准决策的复杂任务,仅靠语言能力是不够的——模型必须具备逻辑思考、经验调用和信息整合以得出结论并采取行动的能力。为提升推理能力,提示(prompting)和微调(fine-tuning)等方法已被广泛探索。尽管这些方法显著提升了推理表现,但其对模型安全的影响仍不明确。本文研究了推理与安全之间的相互作用,揭示了随着推理能力增强而产生的潜在安全风险,暴露了此前被忽视的漏洞。同时,我们探讨了如何利用推理本身来提升安全性,发现了可能的缓解策略。本研究为开发不仅更强大且更可信的模型提供了关键洞见,适用于实际部署场景。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable success across various NLP benchmarks. However, excelling in complex tasks that require nuanced reasoning and precise decision-making demands more than raw language proficiency--LLMs must reason, i.e., think logically, draw from past experiences, and synthesize information to reach conclusions and take action. To enhance reasoning abilities, approaches such as prompting and fine-tuning have been widely explored. While these methods have led to clear improvements in reasoning, their impact on LLM safety remains less understood. In this work, we investigate the interplay between reasoning and safety in LLMs. We highlight the latent safety risks that arise as reasoning capabilities improve, shedding light on previously overlooked vulnerabilities. At the same time, we explore how reasoning itself can be leveraged to enhance safety, uncovering potential mitigation strategies. By examining both the risks and opportunities in reasoning-driven LLM safety, our study provides valuable insights for developing models that are not only more capable but also more trustworthy in real-world deployments.

大模型安全推理能力提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。