针对法律问答中幻觉问题,提出新评测基准与优化方法。
Fine-tuning Large Language Models for Improving Factuality in Legal Question Answering
- 构建法律幻觉评测集LegalHalBench及三项自动评估指标。
- 结合行为克隆与硬样本感知的迭代偏好优化,显著降低幻觉率。
- 适用于法律AI、司法辅助系统研发人员,提升回答可信度。
大语言模型在高风险领域如法律问答中仍面临幻觉(生成错误或虚构信息)的严峻挑战。为此,本文首次提出名为LegalHalBench的基准数据集及三项自动化评估指标,用于衡量模型在法律问答中的常见幻觉现象。进一步提出一种融合行为克隆与新型硬样本感知迭代直接偏好优化(HIPO)的幻觉缓解方法。通过大量真实数据实验验证,该方法在多项指标上实现显著提升,包括新提出的非幻觉法条率(Non-Hallucinated Statute Rate)、法条相关性率(Statute Relevance Rate)、法律主张真实性(Legal Claim Truthfulness),以及传统指标METEOR、BERTScore、ROUGE-L和胜率。
原文摘要 · Abstract (English)
Hallucination, or the generation of incorrect or fabricated information, remains a critical challenge in large language models (LLMs), particularly in high-stake domains such as legal question answering (QA). In order to mitigate the hallucination rate in legal QA, we first introduce a benchmark called LegalHalBench and three automatic metrics to evaluate the common hallucinations when LLMs answer legal questions. We then propose a hallucination mitigation method that integrates behavior cloning and a novel Hard Sample-aware Iterative Direct Preference Optimization (HIPO). We conduct extensive real-data experiments to validate the effectiveness of our approach. Our results demonstrate remarkable improvements in various metrics, including the newly proposed Non-Hallucinated Statute Rate, Statute Relevance Rate, Legal Claim Truthfulness, as well as traditional metrics such as METEOR, BERTScore, ROUGE-L, and win rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。