arXiv:2511.08082cs.AIcs.LG2025-11

为保险精算中的大模型可靠性建立可衡量的监管框架。

Prudential Reliability of Large Language Models in Reinsurance: Governance, Assurance, and Capital Efficiency

  • 构建五支柱框架,将监管要求转化为可执行的生命周期管控
  • 检索增强配置使事实准确率提升至0.90,幻觉减少40%
  • 适合金融监管机构与保险公司评估AI治理有效性

本文提出一种针对再保险领域大型语言模型(LLM)可靠性的审慎评估框架。该框架采用五支柱结构——治理、数据溯源、保障、韧性与监管对齐,将Solvency II、SR 11-7及EIOPA(2025)、NAIC(2023)、IAIS(2024)的监管指引转化为可度量的全生命周期控制机制。通过再保险人工智能可靠性与保障基准(RAIRAB)进行实施,验证嵌入治理的LLM是否满足接地性、透明度与问责性的审慎标准。在六类任务中,检索增强型配置实现0.90的接地准确率,幻觉和解释漂移降低约40%,透明度几乎翻倍。这些机制有效降低风险转移与资本配置中的信息摩擦,表明现有审慎监管原则已能容纳可靠的AI,前提是治理明确、数据可追溯且保障可验证。

原文摘要 · Abstract (English)

This paper develops a prudential framework for assessing the reliability of large language models (LLMs) in reinsurance. A five-pillar architecture--governance, data lineage, assurance, resilience, and regulatory alignment--translates supervisory expectations from Solvency II, SR 11-7, and guidance from EIOPA (2025), NAIC (2023), and IAIS (2024) into measurable lifecycle controls. The framework is implemented through the Reinsurance AI Reliability and Assurance Benchmark (RAIRAB), which evaluates whether governance-embedded LLMs meet prudential standards for grounding, transparency, and accountability. Across six task families, retrieval-grounded configurations achieved higher grounding accuracy (0.90), reduced hallucination and interpretive drift by roughly 40%, and nearly doubled transparency. These mechanisms lower informational frictions in risk transfer and capital allocation, showing that existing prudential doctrines already accommodate reliable AI when governance is explicit, data are traceable, and assurance is verifiable.

大模型治理再保险可靠性评估监管科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。