arXiv:2509.11619cs.CL2025-09EMNLP被引 1

针对法律对话系统幻觉问题,提出检测与缓解框架并验证效果。

HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain

  • 用LLM构建检测系统,基于LLaMA 3.1 8B Instruct模型
  • 最优策略每轮幻觉仅0.4159次,准确率达96.13%
  • 适合需高可靠性的法律、客服等垂域应用

大型语言模型在工业中广泛应用,但仍易产生幻觉,限制其在关键场景的可靠性。本工作聚焦于基于LLaMA 3.1 8B Instruct构建的消费者投诉聊天机器人中的幻觉问题。我们开发了HalluDetect,一个基于LLM的幻觉检测系统,在基准测试中取得68.92%的F1分数,比基线提升22.47%。通过评估五种幻觉缓解架构,发现AgentBot将幻觉降至每轮0.4159次,同时保持最高令牌准确率(96.13%),是最有效的缓解策略。研究结果提供了一套可扩展的幻觉缓解框架,证明优化推理策略能显著提升事实准确性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are widely used in industry but remain prone to hallucinations, limiting their reliability in critical applications. This work addresses hallucination reduction in consumer grievance chatbots built using LLaMA 3.1 8B Instruct, a compact model frequently used in industry. We develop HalluDetect, an LLM-based hallucination detection system that achieves an F1 score of 68.92% outperforming baseline detectors by 22.47%. Benchmarking five hallucination mitigation architectures, we find that out of them, AgentBot minimizes hallucinations to 0.4159 per turn while maintaining the highest token accuracy (96.13%), making it the most effective mitigation strategy. Our findings provide a scalable framework for hallucination mitigation, demonstrating that optimized inference strategies can significantly improve factual accuracy.

幻觉检测法律AILLM优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。