arXiv:2503.03032cs.CL2025-03EMNLP被引 14

用稀疏自编码器检测并减少大模型幻觉,提升查询生成准确性。

SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs

  • 基于稀疏自编码器构建幻觉检测与抑制框架
  • 在三个跨领域数据集上实现最高29.45%的查询准确率提升
  • 适合关注大模型可靠性与查询增强的研究者

尽管大语言模型(LLMs)表现优异,但其常出现幻觉问题,影响关键应用性能。本文提出SAFE框架,通过稀疏自编码器(SAEs)检测并缓解幻觉。虽然幻觉检测与SAE各自已有研究,但二者的协同应用,尤其在幻觉感知的查询增强方面尚未充分探索。为验证SAFE有效性,我们在两个具备SAE支持的模型上,于三个设计用于评估幻觉问题的跨领域数据集上进行测试。实验结果表明,SAFE在所有数据集上均显著提升查询生成准确率并抑制幻觉,准确率最高提升达29.45%。

原文摘要 · Abstract (English)

Despite the state-of-the-art performance of Large Language Models (LLMs), these models often suffer from hallucinations, which can undermine their performance in critical applications. In this work, we propose SAFE, a novel method for detecting and mitigating hallucinations by leveraging Sparse Autoencoders (SAEs). While hallucination detection techniques and SAEs have been explored independently, their synergistic application in a comprehensive system, particularly for hallucination-aware query enrichment, has not been fully investigated. To validate the effectiveness of SAFE, we evaluate it on two models with available SAEs across three diverse cross-domain datasets designed to assess hallucination problems. Empirical results demonstrate that SAFE consistently improves query generation accuracy and mitigates hallucinations across all datasets, achieving accuracy improvements of up to 29.45%.

大模型幻觉稀疏自编码器查询增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。