揭示金融AI系统中确定性缺失的根源,推动可审计的算法可信性。
From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems

- 从系统层面分析三类金融AI的不可复现问题:表格模型、图网络、大模型代理流程。
- 实测发现信用评分解释排名不稳、欺诈检测预测翻转率高、大模型输出因并行产生偏差。
- 提出分层评估框架,用特定指标匹配审计需求,助力监管合规与技术改进。
在受监管的金融场景(如信用风险、欺诈检测、反洗钱)中部署机器学习,暴露出算法可复现性的关键缺陷。早期金融机器学习关注统计问题如回溯测试过拟合,而深度神经网络与生成式AI引入了源于硬件和架构的机械性非确定性。本综述从系统视角分析当前金融AI三大主流模态的可复现性失效:表格模型(事后解释变异性)、图网络(随机采样与时间异步性)、基于大模型的代理工作流(批处理依赖漂移与轨迹偏移)。结合公开金融数据集的首方实验,量化了信用评分中解释排名不稳定性、图神经网络欺诈检测中的预测翻转率,以及张量并行导致的大模型实体抽取输出偏差。提出分层评估框架,将模态特定指标(RBO、D_cos、TDI、PSD)与审计就绪度关联,并报告这些指标间的重叠而非互补关系。
原文摘要 · Abstract (English)
Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities in algorithmic reproducibility. While early financial ML addressed statistical challenges such as backtest overfitting, deep neural networks and Generative AI have introduced mechanical nondeterminism rooted in hardware and architecture. This survey provides a systems perspective on reproducibility failures across three modalities now dominant in financial AI: tabular models (post-hoc explanation variance), graph networks (stochastic sampling and temporal asynchrony), and LLM-based agentic workflows (batch-dependent divergence and trajectory drift). We supplement the literature analysis with first-party experiments on public financial datasets -- quantifying explanation rank instability in credit scoring, prediction flip rates in GNN-based fraud detection, and tensor-parallel-induced output divergence in LLM entity extraction. We propose a layered evaluation framework linking modality-specific metrics (RBO, D_cos, TDI, PSD) to audit readiness, and report where these measures overlap rather than complement one another.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。