arXiv:2509.04455cs.CL2025-09被引 1

首个面向中文保险领域的AI评估基准,填补专业场景评测空白。

INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance

  • 构建覆盖业务、任务、难度的多维评测体系
  • 含3.87万条高质量题例,评估模型回答忠实度与完整性
  • 发现主流大模型在复杂保险场景中仍有明显短板

保险作为全球金融体系的关键组成部分,对AI应用的准确性与可靠性要求极高。现有评测基准多覆盖通用领域,难以体现保险行业的独特性与需求。为此,我们提出INSEva——一个专为评估AI系统在保险领域知识与能力设计的综合性中文基准。该基准包含38,704条来自权威资料的高质量评测样本,涵盖业务领域、任务形式、难度层级及认知-知识维度的多维评估体系,并针对开放问答设计了专属的忠实度与完整度评估方法。通过对8个前沿大语言模型的全面评估,发现尽管通用模型在基础能力上平均得分超80,但在处理复杂真实保险场景时仍存在显著差距。该基准将很快公开。

原文摘要 · Abstract (English)

Insurance, as a critical component of the global financial system, demands high standards of accuracy and reliability in AI applications. While existing benchmarks evaluate AI capabilities across various domains, they often fail to capture the unique characteristics and requirements of the insurance domain. To address this gap, we present INSEva, a comprehensive Chinese benchmark specifically designed for evaluating AI systems' knowledge and capabilities in insurance. INSEva features a multi-dimensional evaluation taxonomy covering business areas, task formats, difficulty levels, and cognitive-knowledge dimension, comprising 38,704 high-quality evaluation examples sourced from authoritative materials. Our benchmark implements tailored evaluation methods for assessing both faithfulness and completeness in open-ended responses. Through extensive evaluation of 8 state-of-the-art Large Language Models (LLMs), we identify significant performance variations across different dimensions. While general LLMs demonstrate basic insurance domain competency with average scores above 80, substantial gaps remain in handling complex, real-world insurance scenarios. The benchmark will be public soon.

大模型评测保险AI中文基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。