arXiv:2509.05691cs.CLcs.AI2025-09Conference of the …被引 1

测试13种嵌入模型对数字信息的敏感度,发现它们普遍不擅长处理细微数值差异。

Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models

  • 用金融场景的合成数据评估模型对数字的编码能力
  • 13个模型在区分2%与20%等数值差异时表现不佳
  • 研究结果对金融、医疗等重数领域有重要警示意义

文本嵌入模型广泛应用于自然语言处理任务,但其性能通常在不涉及数值理解的任务上进行评估。因此,当前嵌入模型能否准确编码文本中的数字内容仍不明确。这一问题至关重要,因为嵌入模型正越来越多地用于金融、医疗等依赖数字信息的领域。例如,公司A市场份额增长2%与增长20%虽同为增长,但含义迥异。本研究旨在检验嵌入模型是否能捕捉此类细微差别。我们使用金融场景的合成数据,评估了13种广泛应用的文本嵌入模型,发现它们普遍难以准确捕捉数值细节。进一步分析揭示了嵌入模型在数值理解上的局限性,为未来提升嵌入模型在数值内容处理方面的能力提供了方向。

原文摘要 · Abstract (English)

Text embedding models are widely used in natural language processing applications. However, their capability is often benchmarked on tasks that do not require understanding nuanced numerical information in text. As a result, it remains unclear whether current embedding models can precisely encode numerical content, such as numbers, into embeddings. This question is critical because embedding models are increasingly applied in domains where numbers matter, such as finance and healthcare. For example, Company X's market share grew by 2\% should be interpreted very differently from Company X's market share grew by 20\%, even though both indicate growth in market share. This study aims to examine whether text embedding models can capture such nuances. Using synthetic data in a financial context, we evaluate 13 widely used text embedding models and find that they generally struggle to capture numerical details accurately. Our further analyses provide deeper insights into embedding numeracy, informing future research to strengthen embedding model-based NLP systems with improved capacity for handling numerical content.

文本嵌入数值理解金融NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。