首个日语金融文本嵌入评估基准,填补领域空白
JFinTEB: Japanese Financial Text Embedding Benchmark

- 构建涵盖检索与分类任务的综合评估框架
- 覆盖情感分析、文档分类等多类真实金融场景
- 适合日语金融文本研究者与产业应用开发者
我们提出JFinTEB,首个专为日语金融文本嵌入设计的综合性评估基准。现有嵌入基准对日语金融文本的语言与领域特异性覆盖不足。本基准包含多样化的任务类别,涵盖检索与分类任务,反映真实的金融文本处理场景。检索任务利用指令遵循数据集和金融文本生成查询,分类任务包括情感分析、文档分类及基于经济调查数据的领域特定分类挑战。我们在多种嵌入模型上进行了广泛评估,涵盖不同规模的日语专用模型、多语言模型及商用嵌入服务。我们公开发布JFinTEB数据集与评估框架(https://github.com/retarfi/JFinTEB),推动后续研究并为日本金融文本挖掘社区提供标准化评估协议。本工作填补了日语金融文本处理资源的关键空白,为领域嵌入研究奠定基础。
原文摘要 · Abstract (English)
We introduce JFinTEB, the first comprehensive benchmark specifically designed for evaluating Japanese financial text embeddings. Existing embedding benchmarks provide limited coverage of language-specific and domain-specific aspects found in Japanese financial texts. Our benchmark encompasses diverse task categories including retrieval and classification tasks that reflect realistic and well-defined financial text processing scenarios. The retrieval tasks leverage instruction-following datasets and financial text generation queries, while classification tasks cover sentiment analysis, document categorization, and domain-specific classification challenges derived from economic survey data. We conduct extensive evaluations across a wide range of embedding models, including Japanese-specific models of various sizes, multilingual models, and commercial embedding services. We publicly release JFinTEB datasets and evaluation framework at https://github.com/retarfi/JFinTEB to facilitate future research and provide a standardized evaluation protocol for the Japanese financial text mining community. This work addresses a critical gap in Japanese financial text processing resources and establishes a foundation for advancing domain-specific embedding research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。