首个韩语金融大模型评估平台,推动韩语金融AI发展
Won: Establishing Best Practices for Korean Financial NLP
- 建立8周运行的韩语金融大模型开放评测平台
- 覆盖5类选择题与1类开放问答,共评估1119次提交
- 发布8万条指令数据集,公开顶尖模型训练策略
本文首次推出面向韩语金融领域的大型语言模型开放评测基准。该平台运行约八周,对1119个提交方案在封闭测试集上进行了评估,涵盖五类多选题任务:金融与会计、股价预测、国内公司分析、金融市场和金融代理任务,以及一类开放问答任务。基于评估结果,我们发布了包含8万条样本的开源指令数据集,并总结了高性能模型普遍采用的训练策略。最后,我们推出了完全开源透明的Won模型,其构建遵循上述最佳实践。我们希望这些贡献能推动韩语及其他语言金融大模型的更好、更安全的发展。
原文摘要 · Abstract (English)
In this work, we present the first open leaderboard for evaluating Korean large language models focused on finance. Operated for about eight weeks, the leaderboard evaluated 1,119 submissions on a closed benchmark covering five MCQA categories: finance and accounting, stock price prediction, domestic company analysis, financial markets, and financial agent tasks and one open-ended qa task. Building on insights from these evaluations, we release an open instruction dataset of 80k instances and summarize widely used training strategies observed among top-performing models. Finally, we introduce Won, a fully open and transparent LLM built using these best practices. We hope our contributions help advance the development of better and safer financial LLMs for Korean and other languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。