评测日语金融文本理解能力,发现大模型仍难应对隐含承诺与术语嵌套。
Ebisu: Benchmarking Large Language Models in Japanese Finance
- 构建日语金融语境下的两个专家标注任务,评估隐含承诺与术语层级提取。
- 顶尖模型在两项任务上表现不佳,规模扩大对性能提升有限。
- 专为日语金融特点设计,适合研究跨语言、跨文化NLP的学者使用。
日语金融文本具有黏着性、后置结构、混合书写系统及高语境沟通特征,依赖间接表达和隐含承诺,对大模型构成严峻挑战。本文提出Ebisu,一个面向日语金融语言理解的基准测试,包含两个基于语言与文化背景、由专家标注的任务:JF-ICR评估面向投资者问答中隐含承诺与拒绝识别能力;JF-TE评估专业披露文本中嵌套金融术语的层级抽取与排序。我们评测了涵盖通用、日语优化及金融领域的大规模开源与专有模型。结果表明,即使最先进的系统在两项任务上仍表现欠佳。模型规模扩大仅带来有限改进,语言与领域适配也未稳定提升性能,凸显显著差距。Ebisu为推动语言文化融合的金融NLP研究提供聚焦基准。所有数据集与评估脚本均已公开。
原文摘要 · Abstract (English)
Japanese finance combines agglutinative, head-final linguistic structure, mixed writing systems, and high-context communication norms that rely on indirect expression and implicit commitment, posing a substantial challenge for LLMs. We introduce Ebisu, a benchmark for native Japanese financial language understanding, comprising two linguistically and culturally grounded, expert-annotated tasks: JF-ICR, which evaluates implicit commitment and refusal recognition in investor-facing Q&A, and JF-TE, which assesses hierarchical extraction and ranking of nested financial terminology from professional disclosures. We evaluate a diverse set of open-source and proprietary LLMs spanning general-purpose, Japanese-adapted, and financial models. Results show that even state-of-the-art systems struggle on both tasks. While increased model scale yields limited improvements, language- and domain-specific adaptation does not reliably improve performance, leaving substantial gaps unresolved. Ebisu provides a focused benchmark for advancing linguistically and culturally grounded financial NLP. All datasets and evaluation scripts are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。