评测大模型在精算领域的专业能力,涵盖知识、案例推理和工具使用。
INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models

- 构建包含1.2万道题的精算评测基准,分知识、案例、实践三部分。
- 顶尖大模型在标准化知识上表现良好,但在案例推理和工具使用中差距明显。
- 适合研究精算AI或需要可靠专业助手的金融从业者使用。
大语言模型在金融推理方面展现出强大潜力,但现有评测多将领域知识、数值推理、长上下文理解与工具使用分开评估,难以反映真实专业工作流程所需的可审计、上下文相关且可执行决策。本文提出 extbf{INS-ActBench},一个全面评估大模型精算专业能力的基准。该基准包含来自16个精算协会公开考试和样题的12,050组问答对,分为三个子集: extbf{INS-Act-Know}(标准化精算知识)、 extbf{INS-Act-Case}(长上下文保险案例推理)和 extbf{INS-Act-Practice}(需验证数值输出的电子表格与R代码任务)。在九个代表性大模型和人类精算专家上的实验表明,前沿模型在标准化知识上表现优异,但在案例推理、工具驱动流程及受司法管辖区影响的实务操作中仍显著落后。INS-ActBench为开发可信赖的专业精算大模型提供了可复现的基础。代码已开源:https://github.com/FDU-INS/INS-ActBench。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This limits their ability to assess realistic professional workflows that require auditable, context-grounded, and tool-executable decisions. We introduce \textbf{INS-ActBench}, a comprehensive benchmark for evaluating professional actuarial capability in LLMs. INS-ActBench contains 12,050 Q\&A pairs from public exams and sample questions released by 16 actuarial associations. It covers three subsets: \textbf{INS-Act-Know} for standardized actuarial knowledge, \textbf{INS-Act-Case} for long-context insurance case reasoning, and \textbf{INS-Act-Practice} for spreadsheet and R-code tasks with verifiable numerical outputs. Experiments on nine representative LLMs and human actuarial experts reveal a clear capability boundary: frontier LLMs perform strongly on standardized knowledge, but remain much weaker in case reasoning, tool-based workflows, and jurisdiction-sensitive practice. INS-ActBench provides a reproducible foundation for developing actuarial LLMs toward reliable professional assistance. The code is available at https://github.com/FDU-INS/INS-ActBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。