arXiv:2602.13213cs.AIcs.HC2026-02

用对抗自检机制提升保险核保AI的可靠性,让机器先自纠再交给人审。

Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique

  • 设计双代理系统,主代理提方案,批判代理挑错,形成内部制衡。
  • 实测将AI幻觉率从11.3%降至3.8%,决策准确率提升至96%。
  • 适合金融、保险等强监管领域,强调人类最终决策权。

商业保险核保是耗时的人工流程,需手动审查大量文件以评估风险并定价。尽管AI能显著提升效率,现有方案缺乏全面推理与内生可靠性机制,在高风险、受监管环境中难以实现完全自动化。本研究提出一种决策负向、人机协同的智能体系统,引入对抗式自检机制作为受控安全架构,用于监管型核保流程。在此系统中,批评者智能体在主智能体提交建议前对其结论进行挑战,构建内部监督机制,填补了高风险场景下AI安全的关键空白。同时,研究建立了一套形式化失败模式分类体系,为决策负向智能体的潜在错误提供结构化识别与管理框架。基于500个专家验证的核保案例实验表明,对抗自检机制将AI幻觉率从11.3%降低至3.8%,决策准确率从92%提升至96%。该框架通过设计确保人类对所有关键决策拥有绝对控制权。结果表明,对抗自检支持在受监管领域更安全地部署AI,为人类监督不可或缺的场景提供了负责任的整合范式。

原文摘要 · Abstract (English)

Commercial insurance underwriting is a labor-intensive process that requires manual review of extensive documentation to assess risk and determine policy pricing. While AI offers substantial efficiency improvements, existing solutions lack comprehensive reasoning and internal mechanisms to ensure reliability in regulated, high-stakes environments. Full automation remains impractical and inadvisable when human judgment and accountability are critical. This study presents a decision-negative, human-in-the-loop agentic system that incorporates an adversarial self-critique mechanism as a bounded safety architecture for regulated underwriting workflows. In this system, a critic agent challenges the primary agent's conclusions prior to submitting recommendations to human reviewers. This internal system of checks and balances addresses a critical gap in AI safety for regulated workflows. Additionally, the research develops a formal taxonomy of failure modes to characterize potential errors by decision-negative agents. This taxonomy provides a structured framework for risk identification and management in high-stakes applications. Experimental evaluation using 500 expert-validated underwriting cases demonstrates that the adversarial critique mechanism reduces AI hallucination rates from 11.3% to 3.8% and increases decision accuracy from 92% to 96%. At the same time, the framework enforces strict human authority over all binding decisions by design. These findings indicate that adversarial self-critique supports safer AI deployment in regulated domains and offers a model for responsible integration where human oversight is indispensable.

智能体对抗学习保险科技可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。