arXiv:2409.03257cs.CLcs.AI2024-09NAACL被引 6

跟踪11个月韩语大模型进展,揭示性能提升与排名变化规律

Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard

  • 分析1769个韩语模型,持续追踪11个月性能演变
  • 发现模型规模对不同评测任务表现的相关性随时间变化
  • 揭示排行榜排名模式动态演进,适合关注韩语AI发展的研究者

本文通过为期十一个月的纵向研究,弥补了以往关于Open Ko-LLM Leaderboard研究仅限五个月观察期的局限。通过对1,769个模型的分析,探讨三个核心问题:(1)在不同任务上提升韩语大语言模型性能所面临的长期挑战;(2)模型规模如何影响跨基准任务的表现相关性;(3)排行榜排名模式随时间的变化趋势。研究揭示了韩语大模型发展的持续进步及评估框架的动态演化,为理解模型演进提供了更全面的视角。

原文摘要 · Abstract (English)

This paper conducts a longitudinal study over eleven months to address the limitations of prior research on the Open Ko-LLM Leaderboard, which have relied on empirical studies with restricted observation periods of only five months. By extending the analysis duration, we aim to provide a more comprehensive understanding of the progression in developing Korean large language models (LLMs). Our study is guided by three primary research questions: (1) What are the specific challenges in improving LLM performance across diverse tasks on the Open Ko-LLM Leaderboard over time? (2) How does model size impact task performance correlations across various benchmarks? (3) How have the patterns in leaderboard rankings shifted over time on the Open Ko-LLM Leaderboard?. By analyzing 1,769 models over this period, our research offers a comprehensive examination of the ongoing advancements in LLMs and the evolving nature of evaluation frameworks.

韩语大模型纵向研究排行榜分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。