arXiv:2603.08267cs.AIcs.CE2026-03

通过跨模型引导,大幅降低金融大模型偏见检测的计算成本。

Towards a more efficient bias detection in financial language models

  • 利用不同模型输出的共性特征,提前定位偏见输入
  • 在125,000对样本中,识别出0.58%至6.05%的偏见行为
  • 仅用20%输入对,即可发现73%的模型偏见,适合持续训练场景

金融语言模型中的偏见是其实际应用的主要障碍。现有检测方法依赖大规模语料的穷举变异与成对预测分析,虽有效但计算开销巨大,尤其在大模型持续训练中难以为继。本文对五种金融语言模型进行了大规模偏见研究,基于约1.7万条真实财经新闻,构造超过12.5万对原始-变异样本,分析其在受保护属性上的偏见倾向。结果表明,所有模型在原子(0.58%-6.05%)与交叉(0.75%-5.97%)设置下均存在偏见。更关键的是,各模型在偏见揭示输入上表现出一致模式,支持跨模型引导检测,显著提升效率。例如,仅使用DistilRoBERTa指导生成的20%输入对,即可发现FinMA模型73%的偏见行为。

原文摘要 · Abstract (English)

Bias in financial language models constitutes a major obstacle to their adoption in real-world applications. Detecting such bias is challenging, as it requires identifying inputs whose predictions change when varying properties unrelated to the decision, such as demographic attributes. Existing approaches typically rely on exhaustive mutation and pairwise prediction analysis over large corpora, which is effective but computationally expensive-particularly for large language models and can become impractical in continuous retraining and releasing processes. Aiming at reducing this cost, we conduct a large-scale study of bias in five financial language models, examining similarities in their bias tendencies across protected attributes and exploring cross-model-guided bias detection to identify bias-revealing inputs earlier. Our study uses approximately 17k real financial news sentences, mutated to construct over 125k original-mutant pairs. Results show that all models exhibit bias under both atomic (0.58\%-6.05\%) and intersectional (0.75\%-5.97\%) settings. Moreover, we observe consistent patterns in bias-revealing inputs across models, enabling substantial reuse and cost reduction in bias detection. For example, up to 73\% of FinMA's biased behaviours can be uncovered using only 20\% of the input pairs when guided by properties derived from DistilRoBERTa outputs.

偏见检测金融大模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。