arXiv:2601.11000cs.CLcs.AI2026-01ACL被引 2

个性化大模型会因用户历史产生事实幻觉,该研究提出方法缓解此问题。

When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs

  • 在推理阶段引入轻量级修正机制,分离个人化与事实表征
  • 在多个模型上验证,事实准确率显著提升且个性化能力保持
  • 构建首个联合评估个性化与事实问答的基准测试PFQABench

个性化大语言模型(LLMs)通过适应用户行为提升满意度,但可能导致事实推理偏差。我们发现,当面对事实性问题时,个性化模型会生成与用户历史一致而非客观真实的答案,造成由个人化引发的幻觉,损害事实可靠性并可能传播错误认知,根源在于个人化与事实表征的表示纠缠。为此,我们提出事实保真型个性化引导(FPPS),一种轻量级推理阶段方法,可有效缓解个人化导致的事实扭曲,同时保留个性化行为。我们还构建了首个用于联合评估个性化与事实问答的基准测试PFQABench。在多个主流模型及个性化方法上的实验表明,FPPS显著提升事实准确性,同时维持个性化性能。

原文摘要 · Abstract (English)

Personalized large language models (LLMs) adapt model behavior to individual users to enhance user satisfaction, yet personalization can inadvertently distort factual reasoning. We show that when personalized LLMs face factual queries, there exists a phenomenon where the model generates answers aligned with a user's prior history rather than the objective truth, resulting in personalization-induced hallucinations that degrade factual reliability and may propagate incorrect beliefs, due to representational entanglement between personalization and factual representations. To address this issue, we propose Factuality-Preserving Personalized Steering (FPPS), a lightweight inference-time approach that mitigates personalization-induced factual distortions while preserving personalized behavior. We further introduce PFQABench, the first benchmark designed to jointly evaluate factual and personalized question answering under personalization. Experiments across multiple LLM backbones and personalization methods show that FPPS substantially improves factual accuracy while maintaining personalized performance.

大模型个性化幻觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。