为金融模型输出的可信度设计新评估标准,聚焦实际业务中的可辩护性。
Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable
- 构建覆盖七维度的金融工作流输出评估体系,从真实文档中检验可靠性
- 前沿闭源模型平均得分4.3,开源基线仅3.15,差距主要在信息检索与合成环节
- 强调决策有用性维度最能区分模型优劣,适合监管与投行场景使用
在资本市场的实际工作中,关键问题不在于大模型能否生成流畅文本,而在于该文本是否具备可辩护性——即能在对家或监管机构面前站得住脚。现有方法仅覆盖部分需求:通用问答基准关注表面准确性,金融专用基准(FinanceBench、FinQA、ConvFinQA)虽推进文档与数值问答,但评估层级停留在问答对,而非从业者实际提交的工作流输出。本文提出资本市场大模型可靠性评分(CM-LRS),在工作流输出层面评估七个维度:事实准确性、证据可追溯性、数值一致性、流程完整性、来源规范性、决策有用性及可审查性/可审计性。每项0-5分,依据受监管环境下评审人员常用信号制定评分标准,总分可按工作流调整。我们在五类典型任务(如债务承销条款提取、并购可比分析、首次公开募股条款提取)上测试,基于美国证监会EDGAR公开文件、英国收购公告及虚构补充数据,由四位独立评审员(来自三种模型家族)评估四款模型。结果发现:第一,前沿闭源模型在四人平均分上高度集中(0.22分差距),其中Sonnet 4.6得4.31,Opus 4.7得4.30,GPT-5.5得4.09;所有评审均将开源基线Llama 3.3 70B(3.15)排末位。第二,性能差距集中在检索(2.23分)和合成(2.15分)环节,而非条款提取(0.84分)。第三,决策有用性维度在发行人画像任务中展现最大跨模型差异(4.0分跨度),且模型间评分相关性高(均值r=0.52)。可解释性易得,可辩护性才是门槛。
原文摘要 · Abstract (English)
In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft, but whether the draft is bankable: defensible in front of a counter-party or a regulator, with the documents in hand. Existing methods address parts of that gap: open-domain QA benchmarks reward surface accuracy, and finance benchmarks (FinanceBench, FinQA, ConvFinQA) advance document-grounded and numerical QA but evaluate at the question-answer layer rather than the workflow outputs practitioners defend. We introduce CM-LRS, a Capital Markets LLM Reliability Score, evaluating outputs at the workflow-output layer across seven dimensions: factual accuracy, evidence traceability, numerical consistency, workflow completeness, source discipline, decision usefulness, and reviewability/auditability. Each is scored 0-5 against a rubric anchored on signals reviewers in regulated settings use; the aggregate is tunable to the workflow. We demonstrate CM-LRS on five workflows (DCM transaction-terms extraction, precedent retrieval, issuer profile synthesis, M&A transaction-comparable reasoning, ECM transaction-terms extraction) over public SEC EDGAR filings, a public UK takeover release, and fictional synthetic supplements, scoring four models against four independent LLM judges spanning three model families. Three findings. First, the frontier closed-source models cluster within 0.22 points on four-judge averaged CM-LRS (Sonnet 4.6 = 4.31, Opus 4.7 = 4.30, GPT-5.5 = 4.09); all four judges place the open-weights baseline (Llama 3.3 70B = 3.15) last. Second, that gap concentrates on retrieval (2.23) and synthesis (2.15), not extraction (0.84). Third, Decision Usefulness shows the widest cross-model dispersion of any dimension (4.0 points on issuer profiling) and top-tier inter-judge agreement (mean r = 0.52). Plausibility is cheap. Bankability is the bar.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。