arXiv:2607.01951physics.soc-phcs.AI2026-07

大模型面对科学质疑时,不退缩反而更坚定,但效果因模型而异。

Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

论文配图:Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism
图 1 · 摘自论文原文
  • 测试三种模型在气候、疫苗、进化领域的应答策略
  • 发现模型有反应性坚持、表面缓和、不回应三种不同表现
  • 能抵抗质疑的模型未必真懂,可能只是没感知到压力

大型语言模型在应对科学争议问题时,常被担忧会因用户质疑而放弃共识立场,造成虚假平衡。本文在三个指令微调模型(Llama-3.1-8B、Qwen2.5-7B、Mistral-7B)上,针对气候、疫苗、进化三大科学领域,在单轮与多轮对话中进行行为测量、线性探测与激活修补分析。结果未发现顺从性退让,而是观察到三类策略:反应性坚持(共识主张增强,如 Llama)、表面缓和(语气软化但立场不变,如 Qwen)、非响应(如 Mistral)。成对判断确认反应性变化为立场而非风格改变(63.6%,p=0.007),分解分析显示共识主张上升是主因(β=+0.042/剂量,p<1e-77)。线性探测显示差异集中于中间层——Llama 和 Qwen 实现完美分离,而 Mistral 仅 72%,置信区间无重叠,表明其根本未线性表征质疑信号。关键的是,这种鲁棒性不可迁移:跨领域衰减,且在疫苗安全领域甚至逆转,反谣言能力在质疑下减弱。研究提出四类鲁棒性分类框架,强调仅靠行为评估无法区分模型是主动抵抗还是误判信号。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantically retreat from established consensus when a user signals doubt -- drifting toward a false balance that treats settled science as one view among several. We test this across three open instruction-tuned models (Llama-3.1-8B, Qwen2.5-7B, Mistral-7B), three consensus-science domains (climate, vaccines, evolution), and single- and multi-turn settings, combining behavioral measurement with linear probing and activation patching. We do not observe sycophantic retreat. Instead, models show three distinct policies under the same skeptical pressure: reactive assertion, where consensus assertion increases rather than decreases (Llama); surface hedging, where tone softens while the position holds (Qwen); and non-response (Mistral). Pairwise judgments confirm the reactive shift is stance, not style (63.6%, p=.007), and a decomposition identifies increased consensus assertion, not false balance, as its driver (beta=+0.042 per dose, p<1e-77). Linear probes localize the divergence to middle layers -- perfect separation in Llama and Qwen versus 72% in Mistral, with non-overlapping confidence intervals -- indicating the non-responsive model does not linearly represent the skepticism signal at all. Crucially, this robustness does not transfer: it attenuates across domains and, in the safety-critical vaccine domain, can reverse, with myth-rebuttal weakening under skeptical pressure. We synthesize these into a four-way taxonomy separating active from accidental robustness, and argue that behavioral evaluation alone cannot distinguish a model that resists skepticism because it understands the signal from one that only appears to resist because it fails to perceive it.

大模型科学共识鲁棒性对抗性压力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。