arXiv:2608.20661cs.AIcs.CE2026-08

为金融大模型设计可审计的推理框架,让每条答案都有来源可查。

Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance

  • 构建六阶段知识系统,用上下文感知检索追踪证据关系与可信度。
  • 在145个金融问题上,正确率相近但可审计性显著提升。
  • 适合需要合规溯源的金融、审计等高信任场景使用。

企业在金融领域采用大语言模型的障碍不在于语言流畅性,而在于信任:在财务规划与分析(FP&A)等受监管流程中,只有能追溯到权威来源并事后可审计的答案才可用。本文主张,企业金融领域的检索增强生成应以可审计性作为核心评价指标之一,并提出知识驱动分析框架(KDAF)。该框架通过六个迭代阶段构建本体驱动的知识系统,利用上下文感知相关性传播(CARP)检索证据,使每个检索事实均携带关系类型、置信度和来源链。在FinanceBench(145个问题)上的评估表明:零上下文推理正确率仅4.1%,而检索增强方法达10%-12%;各检索方法在答案正确率上无显著差异(如KDAF vs BM25:-0.007,95%CI [-0.021, 0.000]),但可审计性差异明显——KDAF在引用可追溯性F1上达0.515,优于无根基遍历+0.027(CI [0.006, 0.050]),优于BM25 +0.052(CI [0.024, 0.083]),区间均不含零。图结构检索无法引入问题主体外的证据(0/426项),而词法基线分别为16.8%和20.2%;所有选中项均具备完整出处链。因此,本研究认为,在金融场景中,可审计性才是本体增强检索值得投入的核心依据。

原文摘要 · Abstract (English)

Enterprise adoption of large language models in finance is constrained less by fluency than by trust: in Financial Planning and Analysis (FP&A) and other regulated workflows, an answer is usable only if it is traceable to authoritative sources and auditable after the fact. This paper argues that retrieval-augmented generation for enterprise finance should be evaluated on auditability alongside accuracy, and presents the Knowledge-Driven Analytics Framework (KDAF), which builds ontology-driven knowledge systems through six iterative stages and retrieves evidence via Context-Aware Relevance Propagation (CARP), so that every retrieved fact carries its relationship type, confidence, and source lineage. An evaluation on FinanceBench (145 questions) compares KDAF against zero-context inference, BM25, concept-weighted lexical retrieval, and ungrounded graph traversal. First, retrieval is necessary: zero-context inference reaches 4.1% correctness against 10-12% for retrieval-augmented conditions. Second, on answer correctness the retrieval conditions are statistically indistinguishable (KDAF vs BM25: -0.007, 95% CI [-0.021, 0.000]), so accuracy alone does not justify structured retrieval here -- a negative result we report explicitly. Third, on auditability the ordering reverses: KDAF attains the highest citation traceability F1 (0.515), exceeding ungrounded traversal by +0.027 (CI [0.006, 0.050]) and BM25 by +0.052 (CI [0.024, 0.083]), intervals excluding zero. Graph-structured retrieval also admits no evidence from outside the question subject entity (0 of 426 items, against 16.8% and 20.2% for lexical baselines), and every selected item resolves to a complete provenance chain. We argue that auditability, not accuracy, is the axis on which ontology-grounded retrieval earns its cost.

大模型审计金融AI可解释性知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。