arXiv:2410.09623cs.CL2024-10中稿 · EMNLP被引 4

用检索增强生成提升保险问答准确率,但仍有5%-13%错误风险

Quebec Automobile Insurance Question-Answering With Retrieval-Augmented Generation

  • 构建魁北克车险专家语料库与公众问题集,用于评估问答系统
  • 相比纯大模型,引入知识检索后回答准确率显著提升
  • 尽管效果改善,仍存在5%至13%的错误陈述,不适合直接大规模应用

大型语言模型在多种下游任务中表现优异,检索增强生成(RAG)架构已被证明可提升法律问答性能。然而,在保险类问答这一特定法律文档领域,相关应用仍有限。本文构建了两个数据集:魁北克车险专家参考语料库,以及82个面向普通民众的车险问题及专家解答。研究利用这两个语料库,对当前最先进的GPT4-o模型进行自动与人工评估,以回答魁北克车险问题。结果表明,使用专家参考语料库后,自动与人工评估指标均有所提升。然而,研究也揭示大模型问答在关键领域仍不可靠:5%至13%的回答包含可能引发客户误解的错误陈述。

原文摘要 · Abstract (English)

Large Language Models (LLMs) perform outstandingly in various downstream tasks, and the use of the Retrieval-Augmented Generation (RAG) architecture has been shown to improve performance for legal question answering (Nuruzzaman and Hussain, 2020; Louis et al., 2024). However, there are limited applications in insurance questions-answering, a specific type of legal document. This paper introduces two corpora: the Quebec Automobile Insurance Expertise Reference Corpus and a set of 82 Expert Answers to Layperson Automobile Insurance Questions. Our study leverages both corpora to automatically and manually assess a GPT4-o, a state-of-the-art LLM, to answer Quebec automobile insurance questions. Our results demonstrate that, on average, using our expertise reference corpus generates better responses on both automatic and manual evaluation metrics. However, they also highlight that LLM QA is unreliable enough for mass utilization in critical areas. Indeed, our results show that between 5% to 13% of answered questions include a false statement that could lead to customer misunderstanding.

保险问答RAG大模型评估知识检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。