arXiv:2410.12835cs.CEcs.CL2024-10中稿 · ACM ICAIF'24被引 8

首个专为荷兰金融任务优化的中文大模型,支持多语言金融数据构建。

A Dutch Financial Large Language Model

  • 基于自动化翻译与处理构建14万样本荷兰金融指令数据集
  • 在5项荷兰及英语金融任务中表现优于现有模型
  • 提供可复用的开源数据构建方法,适合金融NLP研究者

本文提出FinGEITje,首个专为荷兰金融任务设计和优化的大型语言模型。伴随模型发布,我们构建了一个包含超过14万样本的专用荷兰语金融指令微调数据集,采用自动化翻译与数据处理方法生成。开放源代码的数据构建流程可支持其他语言的金融指令数据集创建。为评估模型性能,研究引入首个荷兰语金融评估基准,并提出一种利用大语言模型作为独立评估者的自动化评估方法,显著减少人工干预。实验结果表明,FinGEITje在五项关键荷兰语和英语金融任务中均展现出卓越性能。

原文摘要 · Abstract (English)

This paper presents FinGEITje, the first Dutch financial Large Language Model (LLM) specifically designed and optimized for various financial tasks. Together with the model, we release a specialized Dutch financial instruction tuning dataset with over 140,000 samples, constructed employing an automated translation and data processing method. The open-source data construction method is provided, facilitating the creation of financial instruction datasets in different languages. To evaluate model performance, the study introduces the first Dutch financial evaluation benchmark, along with an automated evaluation method that utilizes an LLM as an independent evaluator, reducing manual intervention in performance evaluation. The experimental results highlight the superior performance of FinGEITje across five critical Dutch and English financial tasks.

金融大模型多语言数据构建自动评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。