arXiv:2511.07794cs.CL2025-11被引 2

首个保险大模型评估基准,系统评测11个主流模型表现

Design, Results and Industry Implications of the World's First Insurance Large Language Model Evaluation Benchmark

  • 构建覆盖5维度54指标的保险领域评估框架
  • 发现通用模型在精算与合规上普遍薄弱
  • 适合保险AI研发与行业选型参考

本文详述了全球首个保险大语言模型评估基准CUFEInse v1.0的设计方法、多维评价体系及底层理念。遵循“量化导向、专家驱动、多验证”原则,该基准涵盖5个核心维度、54个子指标和14,430个高质量问题,涉及保险理论知识、行业理解、安全合规、智能代理应用与逻辑严谨性。基于此,对11个主流大模型进行了全面评估。结果表明,通用模型普遍存在精算能力弱、合规适应差等共性瓶颈;高质量垂直训练虽在保险场景中表现更优,但在业务适配与合规性上仍有不足。评估精准识别出当前大模型在保险精算、核保理赔推理及合规营销文案生成中的常见短板。CUFEInse的建立填补了保险领域专业评估基准空白,为学术研究与产业应用提供权威工具,其构建理念也为垂直领域大模型评估提供重要参考。最后,论文展望了评估基准的迭代方向,提出保险大模型发展应聚焦‘领域适配+推理增强’。

原文摘要 · Abstract (English)

This paper comprehensively elaborates on the construction methodology, multi-dimensional evaluation system, and underlying design philosophy of CUFEInse v1.0. Adhering to the principles of "quantitative-oriented, expert-driven, and multi-validation," the benchmark establishes an evaluation framework covering 5 core dimensions, 54 sub-indicators, and 14,430 high-quality questions, encompassing insurance theoretical knowledge, industry understanding, safety and compliance, intelligent agent application, and logical rigor. Based on this benchmark, a comprehensive evaluation was conducted on 11 mainstream large language models. The evaluation results reveal that general-purpose models suffer from common bottlenecks such as weak actuarial capabilities and inadequate compliance adaptation. High-quality domain-specific training demonstrates significant advantages in insurance vertical scenarios but exhibits shortcomings in business adaptation and compliance. The evaluation also accurately identifies the common bottlenecks of current large models in professional scenarios such as insurance actuarial, underwriting and claim settlement reasoning, and compliant marketing copywriting. The establishment of CUFEInse not only fills the gap in professional evaluation benchmarks for the insurance field, providing academia and industry with a professional, systematic, and authoritative evaluation tool, but also its construction concept and methodology offer important references for the evaluation paradigm of large models in vertical fields, serving as an authoritative reference for academic model optimization and industrial model selection. Finally, the paper looks forward to the future iteration direction of the evaluation benchmark and the core development direction of "domain adaptation + reasoning enhancement" for insurance large models.

大模型评估保险AI垂直领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。