评测大模型在金融风险审查中的决策能力,关注证据与操作匹配度。
FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

- 构建双维度评测框架:固定证据下的操作执行和动态证据控制。
- 33个模型配置下操作级排名不冗余,知识筛选最多导致18.01分失误。
- 模型频繁提问未必提升决策质量,提示需对齐真实工作流程。
将大语言模型用于专业金融审查,不仅需评估通用财务能力,更需检验其能否完成特定审查任务并判断证据是否足以支撑可辩护的决策。现有金融基准多聚焦知识、推理、合规与专业任务,但评估单元常以数据集或任务形式组织,而非部署系统所支持的决策。本文提出中文金融风险评测基准FinRiskAtlas,从两个互补维度评估金融LLM:在固定证据状态下的操作执行,以及在动态审查条件下对证据状态的控制能力。静态基准包含53个任务族共9,742个实例,涵盖42个领域知识类与11个下游审查操作,均由明确评估合约定义。FinRisk-Ask在此基础上,通过回放104条匿名专业轨迹中的680个预操作状态,推理时隐藏未来证据,仅用其构建专家验证的证据目标。在33种模型配置下,操作级评估产生非冗余排序(下游操作间平均斯皮尔曼相关系数0.42),基于知识的筛选可能导致单个操作最高达18.01点的损失。此外,进入询问分支频率更高并不必然提升请求定位或端到端证据获取效果。结果表明,广泛的能力评分无法全面反映模型在专业工作流中的可靠性,亟需以决策和证据状态为单位的评测体系。
原文摘要 · Abstract (English)
Deploying large language models for professional financial review requires more than measuring general financial competence: models must perform the specific review operation required by a workflow and determine whether available evidence is sufficient for a defensible decision. Existing financial benchmarks cover knowledge, reasoning, compliance, and professional tasks, but their evaluation units are often organized around datasets or task formulations rather than the decisions that deployed systems support. We introduce FinRiskAtlas, a Chinese-language benchmark that evaluates financial LLMs along two complementary dimensions: operation execution under fixed evidence states and evidence-state control under evolving review conditions. The static benchmark contains 9,742 instances across 53 task families, including 42 Domain Knowledge families and eleven downstream review operations defined by explicit evaluation contracts. FinRisk-Ask extends this framework through offline replay of 680 pre-action states from 104 de-identified professional trajectories, withholding future evidence during inference and using it only to construct expert-verified evidence targets. Across 33 model configurations, operation-level evaluation yields non-redundant rankings (mean pairwise Spearman correlation 0.42 across downstream operations), and knowledge-based shortlisting can incur up to 18.01 points of regret on individual operations. FinRisk-Ask further shows that entering the Ask branch more frequently does not necessarily improve request targeting or end-to-end evidence acquisition. These results show that broad financial capability scores do not fully capture where models are reliable in professional workflows, motivating evaluation units aligned with the decisions and evidence states that deployed systems must support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。