毒性提示会降低大模型事实准确性,且影响其内部计算机制。
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
- 通过控制提示语的语气和词汇,测试对模型输出的影响。
- 毒性提示使事实准确率下降,不确定性升高,而礼貌提示影响有限。
- 发现毒性会激活特定敏感节点,暴露模型内部脆弱性,适合安全研究者参考。
大型语言模型(LLMs)越来越多地应用于对话场景,用户语气从礼貌到攻击性或有毒不等,但关于有毒语言在语义等价提示中是否会影响事实可靠性仍知之甚少。我们研究了词汇和语气类提示扰动如何影响LLM的事实可靠性。通过在礼貌、随机及三种毒性水平下进行受控提示变异,在ARC-Easy、GSM8K和MMLU三个数据集上评估五种LLMs。结果表明,毒性词汇扰动持续降低事实准确性并增加不确定性,而礼貌表述仅带来微弱且不一致的变化。为探究这些答案不一致性是否对应内部变化,我们对模型激活和影响进行了归因图分析。发现毒性增强会特异性放大对扰动敏感的变体节点,而相对稳定的内核推理节点则保持较高不变性。这些发现将提示语气视为LLM可靠性的重要维度,并提供了行为与机制层面的证据:表面词汇变化可改变事实输出与内部计算过程。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in conversational settings where user tone ranges from polite to adversarial or toxic, yet less is known about whether toxic language in otherwise semantically equivalent prompts can degrade factual reliability. We study how lexical and tone-based prompt perturbations affect the factual reliability of LLMs. Using controlled prompt variations across polite, random, and three toxicity levels, we evaluate five LLMs on ARC-Easy, GSM8K, and MMLU. We find that toxic lexical perturbations consistently reduce factual accuracy and increase uncertainty, while polite phrasing yields limited and inconsistent changes. To examine whether these answer inconsistencies correspond to internal changes, we conduct attribution-graph analyses of model activations and influences. We find that increasing toxicity selectively amplifies perturbation-sensitive variant nodes while relatively stable core reasoning nodes remain more invariant. These findings position prompt tone as a critical dimension of LLM reliability and provide behavioral and mechanistic evidence that surface-level lexical variation can alter factual outputs and internal computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。