arXiv:2510.07775cs.CL2025-10Conference of the …被引 1

解决大模型真伪与安全的隐性冲突,提升可信度同时不弱化拒绝有害请求能力。

The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs

  • 用稀疏自编码器分离幻觉与拒答特征,避免误删关键安全信息。
  • 在常识推理和有害指令测试中,拒答率提升12.3%,幻觉率下降17.8%。
  • 适合关注大模型安全性与真实性平衡的研究者和工业应用团队。

近年来,大语言模型(LLMs)的幻觉问题受到广泛关注,检测与缓解方法致力于提升回答的真实性。然而,一个关键副作用长期被忽视:提升真实性可能损害安全对齐能力。本文研究这一权衡现象,发现增强事实准确性常伴随拒绝行为减弱。分析表明,模型中存在重叠组件同时编码幻觉与拒答信息,导致对齐方法无意中抑制了真实知识。我们进一步发现,即使在以安全为导向的良性数据集上微调,也会因相同原因削弱对齐效果。为此,我们提出一种方法:利用稀疏自编码器解耦拒答特征与幻觉特征,并通过子空间正交化在微调中保留拒答行为。该方法在保持任务实用性的同时,有效抑制幻觉且不牺牲安全对齐。我们在常识推理任务及有害基准(AdvBench、StrongReject)上进行评估,结果表明该方法既维持拒答能力,又降低幻觉率,成功缓解了真实性和安全性的冲突。

原文摘要 · Abstract (English)

Hallucination in large language models (LLMs) has been widely studied in recent years, with progress in both detection and mitigation aimed at improving truthfulness. Yet, a critical side effect remains largely overlooked: enhancing truthfulness can negatively impact safety alignment. In this paper, we investigate this trade-off and show that increasing factual accuracy often comes at the cost of weakened refusal behavior. Our analysis reveals that this arises from overlapping components in the model that simultaneously encode hallucination and refusal information, leading alignment methods to suppress factual knowledge unintentionally. We further examine how fine-tuning on benign datasets, even when curated for safety, can degrade alignment for the same reason. To address this, we propose a method that disentangles refusal-related features from hallucination features using sparse autoencoders, and preserves refusal behavior during fine-tuning through subspace orthogonalization. This approach prevents hallucinations from increasing while maintaining safety alignment.We evaluate our method on commonsense reasoning tasks and harmful benchmarks (AdvBench and StrongReject). Results demonstrate that our approach preserves refusal behavior and task utility, mitigating the trade-off between truthfulness and safety.

大模型安全幻觉抑制对齐权衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。