用精心设计的虚假文本训练模型,提升幻觉检测能力。
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection
- 用高欺骗性虚假文本作负样本,结合课程学习逐步增加难度。
- 在MedHallu等难测集上性能提升最高达24%。
- 零样本下优于更大规模的先进模型,适合幻觉检测研究者。
由于幻觉文本具有高度欺骗性,使大语言模型准确检测幻觉仍面临挑战。我们发现,幻觉样本通常比传统负样本更具欺骗性,因此将这些精心设计的幻觉文本作为负样本,用于DPO对齐过程。方法引入课程学习策略,依据独立事实核查模型的概率分数下降幅度,从较易样本逐步过渡到更难样本,实现稳定递进的学习。实验表明,使用课程DPO和高质量负样本训练的HaluCheck模型,在多个指标上显著提升性能,尤其在MedHallu和HaluEval等困难基准上提升最高达24%。此外,HaluCheck模型在零样本设置下表现稳健,显著优于多个更大的前沿模型。
原文摘要 · Abstract (English)
Aligning large language models (LLMs) to accurately detect hallucinations remains a significant challenge due to the sophisticated nature of hallucinated text. Recognizing that hallucinated samples typically exhibit higher deceptive quality than traditional negative samples, we use these carefully engineered hallucinations as negative examples in the DPO alignment procedure. Our method incorporates a curriculum learning strategy, gradually transitioning the training from easier samples, identified based on the greatest reduction in probability scores from independent fact checking models, to progressively harder ones. This structured difficulty scaling ensures stable and incremental learning. Experimental evaluation demonstrates that our HaluCheck models, trained with curriculum DPO approach and high quality negative samples, significantly improves model performance across various metrics, achieving improvements of upto 24% on difficult benchmarks like MedHallu and HaluEval. Additionally, HaluCheck models demonstrate robustness in zero-shot settings, significantly outperforming larger state-of-the-art models across various benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。