arXiv:2608.11344cs.CYcs.AI2026-08

金融AI决策难审计?关键在可验证性而非能力大小。

Governing Agentic AI in FinTech

论文配图:Governing Agentic AI in FinTech
图 1 · 摘自论文原文
  • 提出'可验证性缺口'概念,衡量审计能力与实际可追溯性的差距
  • 实测显示大模型虽强却更难复现,本地小模型反而更易审计
  • 适合监管机构、金融机构及高风险领域AI治理研究者

金融机构正将重要决策权交予自主型AI系统,但其治理机制尚未深入研究。本文指出,制约治理的核心并非能力,而是可验证性。定义‘可验证性缺口’为审计需求与实际可解释性、可复现性之间的差距,该差距受验证者、证据标准和审计延迟影响。构建多层级治理理论,并通过三组实验验证,涵盖从30亿参数本地模型到商用前沿系统的九个版本。实验1显示,供应商发布的模型会改变历史金融行为,且其控制参数(如温度、top_p、top_k)不可调,随机种子也不暴露;在最严格控制下,本地模型复现320/320次执行,托管模型复现319/320和959/960次。实验2表明,调度架构是潜在政策层,不同配置下无任何执行记录重复,前沿模型虽更常复现自身动作,但其记录质量未优,差异化损失相当。实验3显示,两个确定性信用模型可完美复现当前动作,但无法恢复历史动作。提出可复现性应作为治理画像而非单一数值,实现基于证据的授权:只有当留存证据支持时,授权才具正当性。该框架适用于需可审计性的高风险领域。

原文摘要 · Abstract (English)

Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.

AI治理金融科技可验证性审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。