金融AI代理难以复现决策,新框架确保结果可重复且可信。
Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents
- 构建多维度评估框架,量化代理的轨迹与决策确定性
- 发现准确率与可复现性无关,需独立测量(相关系数-0.11)
- 小模型靠模式匹配达高确定性但低准确率,大模型表现不一
LLM代理在金融监管审计中面临可复现性难题:相同输入下无法稳定输出一致结果。本文提出确定性-忠实性保障框架(DFAH),用于衡量工具使用型代理在金融场景中的轨迹确定性、决策确定性及证据依赖忠实性。在超过4700次代理运行(7个模型、4个提供商、3个金融基准,每个50例,T=0.0)中,决策确定性与任务准确率无显著相关性(r = -0.11,95% CI [-0.49, 0.31],p = 0.63,n = 21配置),表明高准确率未必可复现,高可复现性也未必准确。小模型(7-20B)通过严格模式匹配实现近似完美确定性,但准确率仅20-42%;前沿模型确定性为50-96%,准确率波动明显。无一模型同时达到完全确定性与高准确率,验证了DFAH多维评估的必要性。研究提供三个金融基准(合规优先级、投资组合约束、DataOps异常处理;各50例)及开源压力测试工具。在这些设置下,采用模式优先架构的一线模型达到了审计复现所需确定性水平。
原文摘要 · Abstract (English)
LLM agents struggle with regulatory audit replay: when asked to reproduce a flagged transaction decision with identical inputs, many deployments fail to return consistent results. We introduce the Determinism-Faithfulness Assurance Harness (DFAH), a framework for measuring trajectory determinism, decision determinism, and evidence-conditioned faithfulness in tool-using agents deployed in financial services. Across 4,700+ agentic runs (7 models, 4 providers, 3 financial benchmarks with 50 cases each at T=0.0), we find that decision determinism and task accuracy are not detectably correlated (r = -0.11, 95% CI [-0.49, 0.31], p = 0.63, n = 21 configurations): models can be deterministic without being accurate, and accurate without being deterministic. Because neither metric predicts the other in our sample, both must be measured independently, which is precisely what DFAH provides. Small models (7-20B) achieve near-perfect determinism through rigid pattern matching at the cost of accuracy (20-42%), while frontier models show moderate determinism (50-96%) with variable accuracy. No model achieves both perfect determinism and high accuracy, supporting DFAH's multi-dimensional measurement approach. We provide three financial benchmarks (compliance triage, portfolio constraints, and DataOps exceptions; 50 cases each) together with an open-source stress-test harness. Across these benchmarks and DFAH evaluation settings, Tier 1 models with schema-first architectures achieved determinism levels consistent with audit replay requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。