为保险核保设计真实企业场景的AI代理评测基准
Benchmarking Agents in Insurance Underwriting Environments
- 联合领域专家构建多轮核保评测任务,融入真实业务知识与干扰因素
- 13个前沿模型在该基准上表现普遍下降,最高准确率模型效率最低
- 揭示了现有智能体在专业领域易幻觉、框架脆弱,需组合式检测方法
随着AI代理进入企业应用,其评估需要反映真实运营复杂性的基准。然而,现有基准过度强调开放领域任务,使用狭窄的准确率指标,缺乏真实复杂性。我们提出UNDERWRITE,一个由领域专家主导、基于多轮保险核保设计的基准,充分融入企业实际挑战:专有业务知识、噪声工具接口及不完整模拟用户,要求精细的信息收集。评估13个前沿模型后发现,研究实验室性能与企业可用性之间存在显著差距:最准确模型并非最高效,即使拥有工具访问权限仍会虚构领域知识,且pass^k指标性能下降20%。结果表明,专家参与基准设计对真实评估至关重要;常见智能体框架存在脆弱性,导致性能误报;专业领域幻觉检测需采用组合式方法。本工作为更贴近企业部署需求的基准开发提供了关键洞见。
原文摘要 · Abstract (English)
As AI agents integrate into enterprise applications, their evaluation demands benchmarks that reflect the complexity of real-world operations. Instead, existing benchmarks overemphasize open-domains such as code, use narrow accuracy metrics, and lack authentic complexity. We present UNDERWRITE, an expert-first, multi-turn insurance underwriting benchmark designed in close collaboration with domain experts to capture real-world enterprise challenges. UNDERWRITE introduces critical realism factors often absent in current benchmarks: proprietary business knowledge, noisy tool interfaces, and imperfect simulated users requiring careful information gathering. Evaluating 13 frontier models, we uncover significant gaps between research lab performance and enterprise readiness: the most accurate models are not the most efficient, models hallucinate domain knowledge despite tool access, and pass^k results show a 20% drop in performance. The results from UNDERWRITE demonstrate that expert involvement in benchmark design is essential for realistic agent evaluation, common agentic frameworks exhibit brittleness that skews performance reporting, and hallucination detection in specialized domains demands compositional approaches. Our work provides insights for developing benchmarks that better align with enterprise deployment requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。