测试大模型信念抵抗能力,发现小模型易被说服,自信提示反而加剧脆弱性。
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
- 基于SMCR框架设计多轮说服对话,评估大模型信念稳定性。
- 3B模型首轮说服即82.5%改变信念,平均仅1.1-1.4轮完成转变。
- 自信提示反使模型更易动摇,微调效果因模型而异,部分仍极脆弱。
大语言模型(LLMs)在问答任务中应用日益广泛,但近期研究显示其易受说服并采纳非事实性信念。本文基于源-信息-渠道-接收者(SMCR)传播框架,系统评估六种主流大模型在事实知识、医疗问答与社会偏见三个领域对说服的敏感性。通过多轮交互分析不同说服策略对模型陈述信念稳定性的影响,并考察自报信心评分是否增强抗说服能力。结果表明,最小模型Llama 3.2-3B表现出极端顺从性,82.5%的信念改变发生在第一轮说服回合(平均结束轮次为1.1–1.4)。出乎意料的是,自报信心提示不仅未提升鲁棒性,反而加速信念瓦解。探索性对抗微调实验显示,模型表现差异显著:GPT-4o-mini达到近完全鲁棒性(98.6%),Mistral 7B提升显著(从35.7%到79.3%),但Llama系列模型即便在自身失败案例上微调后,鲁棒性仍低于14%(RQ1)。研究揭示当前鲁棒性干预存在显著模型依赖性,为构建更可信的LLMs提供关键指导。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly employed in various question-answering tasks. However, recent studies showcase that LLMs are susceptible to persuasion and could adopt counterfactual beliefs. We present a systematic evaluation of LLM susceptibility to persuasion under the \emph{Source--Message--Channel--Receiver} (SMCR) communication framework. Across six mainstream Large Language Models (LLMs) and three domains (factual knowledge, medical QA, and social bias), we analyze how different persuasive strategies influence stated belief stability over multiple interaction turns. We further examine whether verbalized confidence prompting (i.e., eliciting self-reported confidence scores) affects resistance to persuasion. Results show that the smallest model (Llama 3.2-3B) exhibits extreme compliance, with 82.5\% of belief changes occurring at the first persuasive turn (average end turn of 1.1--1.4). Contrary to expectations, verbalized confidence prompting \emph{increases} vulnerability by accelerating belief erosion rather than enhancing robustness. Finally, an exploratory study of adversarial fine-tuning reveals highly model-dependent effectiveness: GPT-4o-mini achieves near-complete robustness (98.6\%), and Mistral~7B improves substantially (35.7\% $\rightarrow$ 79.3\%), but Llama models remain highly susceptible ($<$14\% RQ1) even when fine-tuned on their own failure cases. Together, these findings highlight substantial model-dependent limits of current robustness interventions and offer guidance for developing more trustworthy LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。