arXiv:2511.07585cs.LGcs.AI2025-11被引 8

小模型比大模型更稳定,金融AI需防输出漂移

LLM Output Drift: Cross-Provider Validation & Mitigation for Financial Workflows

  • 用固定种子+贪婪解码+金融结构约束,量化模型输出一致性
  • 120B大模型仅12.5%输出一致,小模型可达100%
  • 提出三层次部署分级与双提供商审计,适配金融合规

金融机构在对账、监管报告和客户沟通中部署大语言模型(LLMs),但非确定性输出(输出漂移)损害可审计性和信任。我们对五种模型架构(7B-120B参数)在受监管金融任务中量化漂移,发现显著负相关:小模型(Granite-3-8B、Qwen2.5-7B)在T=0.0时输出一致性达100%,而GPT-OSS-120B仅为12.5%(95%置信区间:3.5–36.0%),且不受配置影响(p<0.0001,Fisher精确检验)。贡献包括:(i) 金融校准的确定性测试框架,结合贪婪解码(T=0.0)、固定种子及SEC 10-K结构感知检索排序;(ii) 针对RAG、JSON、SQL输出的任务特定不变性检查,使用金融校准的材料性阈值(±5%)及SEC引用验证;(iii) 三层次模型分类系统,支持风险适配的部署决策;(iv) 双提供商验证的审计就绪证明系统。评估了五个模型(Qwen2.5-7B via Ollama,Granite-3-8B via IBM watsonx.ai,Llama-3.3-70B,Mistral-Medium-2505,GPT-OSS-120B)在三个受监管金融任务上的表现。480次运行(每条件n=16)显示,结构化任务(如SQL)在T=0.2下仍稳定,而RAG任务漂移率达25–75%,呈现任务依赖敏感性。跨提供商验证证实确定性行为可在本地与云部署间传递。本框架映射至FSB、BIS、CFTC要求,提供合规AI部署的可行路径。

原文摘要 · Abstract (English)

Financial institutions deploy Large Language Models (LLMs) for reconciliations, regulatory reporting, and client communications, but nondeterministic outputs (output drift) undermine auditability and trust. We quantify drift across five model architectures (7B-120B parameters) on regulated financial tasks, revealing a stark inverse relationship: smaller models (Granite-3-8B, Qwen2.5-7B) achieve 100% output consistency at T=0.0, while GPT-OSS-120B exhibits only 12.5% consistency (95% CI: 3.5-36.0%) regardless of configuration (p<0.0001, Fisher's exact test). This finding challenges conventional assumptions that larger models are universally superior for production deployment. Our contributions include: (i) a finance-calibrated deterministic test harness combining greedy decoding (T=0.0), fixed seeds, and SEC 10-K structure-aware retrieval ordering; (ii) task-specific invariant checking for RAG, JSON, and SQL outputs using finance-calibrated materiality thresholds (plus or minus 5%) and SEC citation validation; (iii) a three-tier model classification system enabling risk-appropriate deployment decisions; and (iv) an audit-ready attestation system with dual-provider validation. We evaluated five models (Qwen2.5-7B via Ollama, Granite-3-8B via IBM watsonx.ai, Llama-3.3-70B, Mistral-Medium-2505, and GPT-OSS-120B) across three regulated financial tasks. Across 480 runs (n=16 per condition), structured tasks (SQL) remain stable even at T=0.2, while RAG tasks show drift (25-75%), revealing task-dependent sensitivity. Cross-provider validation confirms deterministic behavior transfers between local and cloud deployments. We map our framework to Financial Stability Board (FSB), Bank for International Settlements (BIS), and Commodity Futures Trading Commission (CFTC) requirements, demonstrating practical pathways for compliance-ready AI deployments.

金融AI输出漂移大模型合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。