arXiv:2509.03871cs.CLcs.AI2025-09综述被引 22

系统梳理大模型推理可信性的五大维度,揭示当前技术的潜力与风险。

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

  • 按时间线梳理思维链技术在可信性五维度的应用
  • 发现先进推理模型仍存在安全、鲁棒性等严重漏洞
  • 适合关注AI安全与大模型可信性的研究者阅读

长思维链(Long-CoT)推理显著提升了大语言模型在语言理解、复杂问题求解和代码生成等任务上的表现。该范式通过生成中间推理步骤,提升了准确率与可解释性。然而,对基于思维链的推理如何影响模型可信性的全面理解仍不充分。本文综述了近期推理模型与思维链技术的研究,聚焦可信推理的五个核心维度:真实性、安全性、鲁棒性、公平性与隐私保护。针对每个维度,按时间顺序提供研究进展的结构化概述,详析方法、发现与局限。最后提出未来研究方向。尽管推理技术有望通过减少幻觉、检测有害内容、提升鲁棒性来增强可信性,但前沿推理模型本身在安全、鲁棒性和隐私方面仍存在相当甚至更严重的脆弱性。本工作旨在为人工智能安全社区提供及时且有价值的参考资源。相关论文列表见 https://github.com/ybwang119/Awesome-reasoning-safety。

原文摘要 · Abstract (English)

The development of Long-CoT reasoning has advanced LLM performance across various tasks, including language understanding, complex problem solving, and code generation. This paradigm enables models to generate intermediate reasoning steps, thereby improving both accuracy and interpretability. However, despite these advancements, a comprehensive understanding of how CoT-based reasoning affects the trustworthiness of language models remains underdeveloped. In this paper, we survey recent work on reasoning models and CoT techniques, focusing on five core dimensions of trustworthy reasoning: truthfulness, safety, robustness, fairness, and privacy. For each aspect, we provide a clear and structured overview of recent studies in chronological order, along with detailed analyses of their methodologies, findings, and limitations. Future research directions are also appended at the end for reference and discussion. Overall, while reasoning techniques hold promise for enhancing model trustworthiness through hallucination mitigation, harmful content detection, and robustness improvement, cutting-edge reasoning models themselves often suffer from comparable or even greater vulnerabilities in safety, robustness, and privacy. By synthesizing these insights, we hope this work serves as a valuable and timely resource for the AI safety community to stay informed on the latest progress in reasoning trustworthiness. A full list of related papers can be found at \href{https://github.com/ybwang119/Awesome-reasoning-safety}{https://github.com/ybwang119/Awesome-reasoning-safety}.

大模型安全思维链可信推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。