arXiv:2605.05482cs.AIcs.CL2026-05ACL

打造银行领域可落地的精准问答模型,兼顾准确与合规。

FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking

  • 用143M tokens数据+智能筛选训练120亿参数模型,提升回答可信度。
  • 拒绝率12%更安全,比基线高近三倍,优于GPT-4.1的过度拒答。
  • 已部署于40多家金融机构,响应快3-5倍,成本降20-50倍。

大型语言模型在各领域快速应用,但银行业因对准确性、合规性及可验证回应的严苛要求,对其采纳仍存阻力。本文提出一个统一、高效的数据驱动框架,用于训练具备领域特性的生成式问答模型,在真实部署条件下优化答案质量、引用溯源与可控拒答能力。首先,构建仅需143M tokens的自动化数据生成流程,结合大模型评分过滤、引用标注与课程学习策略,训练出120亿参数模型,其在引用溯源方面超越GPT-4.1,仅略有引用数量下降。其次,设计校准拒答机制:在22%不可回答样本上训练后,拒答率为12%,显著优于基线模型的4.3%不安全拒答,且避免了GPT-4.1的20.2%过度拒答问题。第三,提出从数据整理到量化部署的全流程方法,系统已在40余家金融机构上线,查询解决率提升7.1个百分点(p < 0.001)。同时,模型响应速度为GPT-4.1的3-5倍,成本降低20-50倍。

原文摘要 · Abstract (English)

Large language models (LLMs) are rapidly being adopted across various domains. However, their adoption in banking industry faces resistance due to demands for high accuracy, regulatory compliance, and the need for verifiable and grounded responses. We present a unified, data-efficient framework for training grounded domain-specific LLMs that optimizes answer quality, citation grounding, and calibrated refusal under real-world deployment constraints. First, we describe a data generation pipeline that combines LLM-as-a-Judge filtering, citation annotation, and curriculum learning with only 143M tokens. The resulting 12B model achieves high answer quality outperforming GPT-4.1 on citation grounding, with a modest citation tradeoff versus the untuned base. Second, we propose a calibrated refusal mechanism: training on 22% unanswerable examples yield a 12% "I don't know" rate, substantially improving over the base model's unsafe 4.3% rate while avoiding GPT-4.1's over-refusal (20.2%). Third, we present an end-to-end methodology spanning from data curation to quantized serving. The system is deployed at 40+ financial institutions, achieving a 7.1 percentage point improvement in query resolution (p < 0.001). Additionally, the model delivers 3-5x faster responses at 20-50x lower cost compared to GPT-4.1.

金融AI问答系统大模型落地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。