arXiv:2502.19127cs.CL2025-02EMNLP被引 4

通过强化模型精准用知识能力,有效减少大模型幻觉问题。

Exploring the Generalizability of Factual Hallucination Mitigation via Enhancing Precise Knowledge Utilization

  • 用偏好优化微调模型,提升其对事实性问题的精准回答能力。
  • 在18.1万条中文事实问答数据上训练,跨21个领域表现稳定提升。
  • 适用多语言、多任务,尤其适合需要高准确性的场景使用。

大语言模型常因无法与客观事实对齐而产生事实性幻觉,难以被察觉且易误导用户。尽管已有后训练方法缓解此问题,但普遍存在泛化能力差和与其他能力存在权衡的缺陷。本文提出通过直接增强模型精准利用知识的能力来解决上述问题,引入PKUE(精确知识利用增强)方法:通过偏好优化,对模型在自生成的事实性简单问答上的响应进行微调。同时构建了FactualBench,一个包含18.1万条中文数据、覆盖21个领域的综合性、高精度事实问答数据集,用于评估与训练。大量实验表明,PKUE显著提升了模型整体表现,在多种形式的事实任务、非事实性通用任务以及不同语言的任务中均实现一致增强。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often struggle to align their responses with objective facts, resulting in the issue of factual hallucinations, which can be difficult to detect and mislead users without relevant knowledge. Although post-training techniques have been employed to mitigate the issue, existing methods usually suffer from poor generalization and trade-offs in other different capabilities. In this paper, we propose to address these by directly augmenting LLM's fundamental ability to precisely leverage its knowledge and introduce PKUE (Precise Knowledge Utilization Enhancement), which fine-tunes the model on self-generated responses to precise and simple factual questions through preference optimization. Furthermore, we construct FactualBench, a comprehensive and precise factual QA dataset containing 181k Chinese data spanning 21 domains, to facilitate both evaluation and training. Extensive experiments demonstrate that PKUE significantly improves LLM overall performance, with consistent enhancement across factual tasks of various forms, general tasks beyond factuality, and tasks in different language.

大模型幻觉抑制知识利用中文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。