arXiv:2603.14463cs.CL2026-03

保险大模型突破领域专精与通用能力的矛盾,实现零幻觉推理。

An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs

  • 构建可验证数据合成系统与渐进式强化学习训练框架。
  • 在3.9万+样本保险基准上表现领先,幻觉率低至0.6%。
  • 适合需要高可靠性的金融保险AI落地场景。

将大语言模型适配至保险等高风险垂直领域面临严峻挑战:必须严格遵守复杂法规与业务逻辑,对幻觉零容忍。现有方法常陷入能力权衡——牺牲通用智能换取领域专长,或过度依赖RAG而缺乏内在推理。为此,我们提出INS-S1保险专用大模型家族,采用端到端对齐范式。方法创新包括:(1) 可验证数据合成系统,构建分层的精算与合规数据集;(2) 渐进式SFT-RL课程框架,融合动态数据退火、经验证推理(RLVR)与AI反馈(RLAIF)。通过优化数据比例与奖励信号,在强化领域约束的同时防止灾难性遗忘。此外,我们发布了目前最全面的保险评估基准INSEva(含39,000+样本)。大量实验表明,INS-S1在领域任务上达到当前最优性能,显著优于DeepSeek-R1和Gemini-2.5-Pro。关键的是,其保持顶级通用能力,幻觉率降至0.6%(HHEM)。结果证明,严谨的领域专业化无需以牺牲通用智能为代价。

原文摘要 · Abstract (English)

Adapting Large Language Models (LLMs) to high-stakes vertical domains like insurance presents a significant challenge: scenarios demand strict adherence to complex regulations and business logic with zero tolerance for hallucinations. Existing approaches often suffer from a Competency Trade-off - sacrificing general intelligence for domain expertise - or rely heavily on RAG without intrinsic reasoning. To bridge this gap, we present INS-S1, an insurance-specific LLM family trained via a novel end-to-end alignment paradigm. Our approach features two methodological innovations: (1) A Verifiable Data Synthesis System that constructs hierarchical datasets for actuarial reasoning and compliance; and (2) A Progressive SFT-RL Curriculum Framework that integrates dynamic data annealing with a synergistic mix of Verified Reasoning (RLVR) and AI Feedback (RLAIF). By optimizing data ratios and reward signals, this framework enforces domain constraints while preventing catastrophic forgetting. Additionally, we release INSEva, the most comprehensive insurance benchmark to date (39k+ samples). Extensive experiments show that INS-S1 achieves SOTA performance on domain tasks, significantly outperforming DeepSeek-R1 and Gemini-2.5-Pro. Crucially, it maintains top-tier general capabilities and achieves a record-low 0.6% hallucination rate (HHEM). Our results demonstrate that rigorous domain specialization can be achieved without compromising general intelligence.

保险AI幻觉控制大模型训练领域专精

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。