arXiv:2507.21132cs.AIcs.CY2025-07被引 6

测试大模型在人生重大决策中的可靠性,发现部分模型会盲目迎合用户。

Can You Trust an LLM with Your Life-Changing Decision? An Investigation into AI High-Stakes Responses

  • 通过压力测试和追问机制评估模型稳定性
  • 顶尖模型靠频繁提问而非直接建议来提升安全性
  • 可操控激活向量实现安全行为的精准控制

大型语言模型(LLMs)正被越来越多地用于人生重大决策咨询,但其缺乏防止自信却错误回应的标准防护机制,易出现迎合用户与过度自信的问题。本文通过三项实验进行探究:(1)多选题评估模型在用户压力下的响应稳定性;(2)基于新安全分类体系与LLM裁判的自由回答分析;(3)通过操纵“高风险”激活向量的机制可解释性实验,调控模型行为。结果表明,部分模型存在迎合倾向,而如o4-mini等模型则表现稳健。表现最优的模型通过频繁提出澄清问题来获得高安全评分,体现谨慎、求证的特质,而非直接给出建议。此外,我们证明可通过激活向量操控直接调节模型的审慎程度,为安全对齐提供新路径。研究强调需建立多维度、精细化的评估基准,以确保大模型在生死攸关决策中可被信赖。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly consulted for high-stakes life advice, yet they lack standard safeguards against providing confident but misguided responses. This creates risks of sycophancy and over-confidence. This paper investigates these failure modes through three experiments: (1) a multiple-choice evaluation to measure model stability against user pressure; (2) a free-response analysis using a novel safety typology and an LLM Judge; and (3) a mechanistic interpretability experiment to steer model behavior by manipulating a "high-stakes" activation vector. Our results show that while some models exhibit sycophancy, others like o4-mini remain robust. Top-performing models achieve high safety scores by frequently asking clarifying questions, a key feature of a safe, inquisitive approach, rather than issuing prescriptive advice. Furthermore, we demonstrate that a model's cautiousness can be directly controlled via activation steering, suggesting a new path for safety alignment. These findings underscore the need for nuanced, multi-faceted benchmarks to ensure LLMs can be trusted with life-changing decisions.

大模型安全人生决策可信AI可控性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。