arXiv:2511.09997cs.CL2025-11被引 1

发现BERTScore在金融文本中对数值差异不敏感,易误判语义

FinNuE: Exposing the Risks of Using BERTScore for Numerical Semantic Evaluation in Finance

  • 构建含受控数值扰动的金融语料诊断集FinNuE
  • BERTScore对关键数值差异区分能力差,相似度评分虚高
  • 适合金融NLP评估研究者关注模型评价缺陷

BERTScore已成为评估自然语言句子语义相似性的常用指标。但我们发现其存在关键局限:对数值变化敏感性低,这在金融领域尤为严重,因数值精度直接影响语义(如2%收益与20%亏损的区别)。我们构建了FinNuE,一个涵盖财报电话会、监管文件、社交媒体和新闻文章的诊断数据集,通过受控数值扰动进行测试。结果表明,BERTScore无法有效区分语义上截然不同的文本对,常赋予财务含义迥异的句子过高相似度评分。研究揭示了基于嵌入的评估指标在金融场景下的根本缺陷,呼吁发展具备数值感知能力的金融NLP评价框架。

原文摘要 · Abstract (English)

BERTScore has become a widely adopted metric for evaluating semantic similarity between natural language sentences. However, we identify a critical limitation: BERTScore exhibits low sensitivity to numerical variation, a significant weakness in finance where numerical precision directly affects meaning (e.g., distinguishing a 2% gain from a 20% loss). We introduce FinNuE, a diagnostic dataset constructed with controlled numerical perturbations across earnings calls, regulatory filings, social media, and news articles. Using FinNuE, demonstrate that BERTScore fails to distinguish semantically critical numerical differences, often assigning high similarity scores to financially divergent text pairs. Our findings reveal fundamental limitations of embedding-based metrics for finance and motivate numerically-aware evaluation frameworks for financial NLP.

金融NLP语义评估数值敏感性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。