arXiv:2507.09601cs.CLcs.AI2025-07中稿 · FinAI@CIKM 2025被引 1

专为金融领域打造的跨语言嵌入模型,提升韩英金融文本理解能力。

NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance

  • 用1.88万组精准语义对微调,融合同义句、难负例与双语翻译
  • 在韩语金融语义匹配任务上比现有模型提升0.22,英语任务提升0.10
  • 公开韩语金融语义匹配数据集KorFinSTS,助力低资源语言研究

通用句子嵌入模型在捕捉金融领域专业语义时表现不佳,尤其在韩语等低资源语言中,受限于领域术语、语义随时间演变及双语词表错配。为此,我们提出NMIXX(Neural eMbeddings for Cross-lingual eXploration of Finance),一套基于18.8K高置信度三元组微调的跨语言嵌入模型,三元组包含领域内同义句、基于语义演变类型生成的难负例以及精确的韩英翻译对。同时,发布KorFinSTS——一个涵盖新闻、披露文件、研究报告和法规的1,921对韩语金融语义相似性基准,可揭示通用基准忽略的细微差异。在七个开源基线对比中,NMIXX的多语言bge-m3版本在English FinSTS上提升Spearman's rho 0.10,KorFinSTS上提升0.22,优于其预训练检查点,并以最大差距超越其他模型,但通用语义任务性能略有下降。分析显示,韩语词汇覆盖更丰富的模型适应性更强,凸显分词器设计在低资源跨语言场景中的关键作用。通过公开模型与基准,为金融领域领域自适应多语言表示学习提供有力工具。

原文摘要 · Abstract (English)

General-purpose sentence embedding models often struggle to capture specialized financial semantics, especially in low-resource languages like Korean, due to domain-specific jargon, temporal meaning shifts, and misaligned bilingual vocabularies. To address these gaps, we introduce NMIXX (Neural eMbeddings for Cross-lingual eXploration of Finance), a suite of cross-lingual embedding models fine-tuned with 18.8K high-confidence triplets that pair in-domain paraphrases, hard negatives derived from a semantic-shift typology, and exact Korean-English translations. Concurrently, we release KorFinSTS, a 1,921-pair Korean financial STS benchmark spanning news, disclosures, research reports, and regulations, designed to expose nuances that general benchmarks miss. When evaluated against seven open-license baselines, NMIXX's multilingual bge-m3 variant achieves Spearman's rho gains of +0.10 on English FinSTS and +0.22 on KorFinSTS, outperforming its pre-adaptation checkpoint and surpassing other models by the largest margin, while revealing a modest trade-off in general STS performance. Our analysis further shows that models with richer Korean token coverage adapt more effectively, underscoring the importance of tokenizer design in low-resource, cross-lingual settings. By making both models and the benchmark publicly available, we provide the community with robust tools for domain-adapted, multilingual representation learning in finance.

跨语言嵌入金融NLP韩语处理语义匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。