金融NLP评测的标签可信度受评分标准影响,需建立评估治理规范。
Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR
- 通过多版本评分标准测试模型标注一致性,发现表述差异导致70%-83%分歧
- 仅精确准确率、宏F1和加权卡帕系数在该数据分布下具有可解释性
- 提出评测结果需经指标可识别性审计,避免排名误导决策
随着大模型在财报电话会、投资者问答、指引和披露文本中的应用日益可信,监督式金融NLP基准已成为模型选型与部署的重要依据。然而,这一判断的基础假设——黄金标签是客观的——在评测标准本身对评分细则、指标选择或聚合策略敏感时便失效。本文在日文财务隐含承诺识别(JF-ICR)任务上进行了系统分析:253个测试样本,4个前沿LLM,5种评分细则,3种温度参数,5种序数指标。结果表明:评分细则表述变化显著影响模型标注,同一批次中不同细则间的一致率在70.0%至83.4%之间波动,尤其集中在+1/0隐含承诺边界附近;在现有类别分布下,单一准确率因多数类主导且近似预测获分而过于宽松,最差类别准确率因稀有类别仅2例而噪声过大,故仅有精确准确率、宏F1与加权κ(weighted kappa)具备可解释性;在此基础上,布拉德利-特里、博达和排序对等方法在可识别指标子集上达成一致,而在全五指标组合下却对最接近的模型对产生分歧。本文贡献不在于新排行榜,而在于为存在黄金标签的监督式金融评测建立评估治理准则。
原文摘要 · Abstract (English)
As LLMs become credible readers of earnings calls, investor-relations Q\&A, guidance, and disclosure language, supervised financial NLP benchmarks increasingly function as decision evidence for model selection and deployment. A hidden assumption is that gold labels make such evidence objective. This assumption breaks down when the benchmark ruler itself is sensitive to rubric wording, metric choice, or aggregation policy. We study this measurement risk on Japanese Financial Implicit-Commitment Recognition (JF-ICR; a pinned 253-item test split x 4 frontier LLMs x 5 rubrics x 3 temperatures x 5 ordinal metrics). Three findings follow. First, rubric wording materially changes model-assigned labels: R2--R3 agreement ranges from 70.0% to 83.4%, with the dominant movement near the +1 / 0 implicit-commitment boundary. This pattern is consistent with a pragmatic-boundary interpretation, but is not a validated linguistic-causality claim because the present rubric variants confound semantics, examples, and verbosity. Second, not every metric remains informative under the JF-ICR class distribution. Within-one accuracy is too easy because near misses receive credit and the majority class dominates; worst-class accuracy is too noisy because the rarest class has only two examples. Exact accuracy, macro-F1, and weighted \k{appa} are therefore the identifiable metrics under our operational rule. Third, ranking claims become more defensible only after this metric-identifiability audit: Bradley--Terry, Borda, and Ranked Pairs agree on the identifiable metric subset, while the full five-metric sweep produces disagreement on the closest pair. The contribution is not a new leaderboard, but a reporting discipline for supervised financial benchmarks whose gold labels exist and whose evaluation ruler still requires governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。