用多智能体自动生成并评估精算题目,验证了自主校验和评测方法的有效性。
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
- 分角色部署四个LLM:出题、造干扰项、独立校验修复、摘要标签
- 8家厂商50个模型测试,开源完整榜单与题目网页可查
- 本地运行的开源模型表现逼近顶尖,且评测方式影响结果排名
我们提出ActuBench,一个面向精算评估任务的多智能体大模型流水线,生成内容与国际精算协会(IAA)教育大纲对齐。流水线通过适配器分离四类角色:一个智能体出题,一个构造干扰项,第三个独立验证前后阶段并驱动有限次修复循环,第四个成本优化辅助代理负责维基百科笔记摘要与主题标注。所有题目、各模型响应及完整排行榜均以可浏览网页形式发布于https://actubench.de/en/,无需克隆仓库即可查看。我们在两个互补基准上评估了来自8家厂商的50个语言模型:100个实证最难的选择题与100个开放题由LLM裁判评分。主要发现:第一,多智能体验证是核心——独立校验者首次即识别多数问题,多数可通过一次修复解决;第二,本地部署的开源模型处于成本-性能帕累托前沿:在消费级硬件运行的Gemma~4模型,以及在Cerebras托管的120B开源模型,在近零成本区域表现突出,后者仅差一题即达榜首;第三,选择题与LLM裁判评价结果显著不同:选择题架构抬高了性能上限,而裁判模式才能在前沿进行有效区分。
原文摘要 · Abstract (English)
We present ActuBench, a multi-agent LLM pipeline for the automated generation and evaluation of advanced actuarial assessment items aligned with the International Actuarial Association (IAA) Education Syllabus. The pipeline separates four LLM roles by adapter: one agent drafts items, one constructs distractors, a third independently verifies both stages and drives bounded one-shot repair loops, and a cost-optimized auxiliary agent handles Wikipedia-note summarization and topic labelling. The items, per-model responses and complete leaderboard are published as a browsable web interface at https://actubench.de/en/, allowing readers and practitioners to inspect individual items without a repository checkout. We evaluate 50 language models from eight providers on two complementary benchmarks -- 100 empirically hardest multiple-choice items and 100 open-ended items scored by an LLM judge -- and report three headline findings. First, multi-agent verification is load-bearing: the independent verifier flags a majority of drafted items on first pass, most of which the one-shot repair loop resolves. Second, locally-hosted open-weights inference sits on the cost-performance Pareto front: a Gemma~4 model running on consumer hardware and a Cerebras-hosted 120B open-weights model dominate the near-zero-cost region, with the latter within one item of the top of the leaderboard. Third, MCQ and LLM-as-Judge rankings differ meaningfully: the MCQ scaffold inflates the performance ceiling, and Judge-mode evaluation is needed to discriminate at the frontier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。