arXiv:2501.10943cs.CLcs.AI2025-01被引 1

构建中文保险领域问答基准,提升大模型专业应用能力

InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models

  • 构建三类保险数据集,覆盖常识、结构化与非结构化场景
  • 微调后模型在保险术语和条款理解上性能显著提升
  • 适合研究保险AI、大模型落地的开发者与研究人员

大语言模型在多个领域取得显著进展,但在中文保险等专业领域的应用仍不充分。保险知识涉及专业术语和多样数据类型,对模型和用户均构成挑战。为此,我们提出InsQABench,一个面向中文保险领域的基准数据集,包含三类任务:保险常识知识、保险结构化数据库和保险非结构化文档,真实反映保险问答场景。我们还提出SQL-ReAct和RAG-ReAct两种方法,分别应对结构化与非结构化数据任务。评估显示,尽管当前大模型在领域术语和复杂条款理解上表现不佳,但在InsQABench上进行微调后性能明显改善。该基准为推动大模型在保险领域的应用奠定了基础,数据与代码已开源:https://github.com/HaileyFamo/InsQABench.git。

原文摘要 · Abstract (English)

The application of large language models (LLMs) has achieved remarkable success in various fields, but their effectiveness in specialized domains like the Chinese insurance industry remains underexplored. The complexity of insurance knowledge, encompassing specialized terminology and diverse data types, poses significant challenges for both models and users. To address this, we introduce InsQABench, a benchmark dataset for the Chinese insurance sector, structured into three categories: Insurance Commonsense Knowledge, Insurance Structured Database, and Insurance Unstructured Documents, reflecting real-world insurance question-answering tasks.We also propose two methods, SQL-ReAct and RAG-ReAct, to tackle challenges in structured and unstructured data tasks. Evaluations show that while LLMs struggle with domain-specific terminology and nuanced clause texts, fine-tuning on InsQABench significantly improves performance. Our benchmark establishes a solid foundation for advancing LLM applications in the insurance domain, with data and code available at https://github.com/HaileyFamo/InsQABench.git.

大模型保险问答系统中文NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。