arXiv:2605.01870cs.CL2026-05

用大模型推理能力蒸馏出高效希腊语大模型,解决小语种资源少难题。

Maistros: A Greek Large Language Model Adapted Through Knowledge Distillation From Large Reasoning Models

论文配图:Maistros: A Greek Large Language Model Adapted Through Knowledge Distillation From Large Reasoning Models
图 1 · 摘自论文原文
  • 从大推理模型蒸馏知识,构建轻量级希腊语模型
  • 创建高质量希腊语问答数据集CulturaQA,覆盖9个任务
  • 适合研究小语种NLP或需部署本地化模型的团队

大型语言模型(LLM)在自然语言处理领域取得显著进展,尤其在多步推理方面表现突出。然而,现有模型在处理超出训练分布的复杂问题时仍易出错,且高参数量大模型推理耗时长,难以在常规设备部署。与此同时,多语言模型研究多集中于高资源语言,对低资源语言支持不足。本文聚焦现代希腊语,针对其问答数据集稀缺的问题,提出:(i) CulturaQA,一个由大推理模型生成并经人工校验的高质量希腊语数据集;(ii) 一种可适配多种语言与任务的轻量级评估框架;(iii) Maistros 8B,一个通过知识蒸馏与Fine-tuning在CulturaQA上训练的开源希腊语大模型;(iv) 在九个由人工构建的希腊语问答数据集上对九个LLM进行全面评估。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have substantially advanced the field of Natural Language Processing (NLP), achieving state-of-the-art performance across a wide range of tasks. These improvements have been attributed, in part, to their emerging reasoning capabilities, which are enabled by large-scale training and increased model capacity. However, existing LLMs can generate erroneous responses when addressing complex queries that fall outside their training distribution, due to limited internal knowledge or the need for multi-step reasoning. To address these limitations, recent work has introduced large reasoning models (LRMs), which incorporate explicit internal reasoning processes to improve response accuracy. Additionally, state-of-the-art LRMs often comprise hundreds of billions of parameters and require several seconds per inference, even on advanced multi-GPU systems. These characteristics limit their practicality for deployment in conventional computing environments. Meanwhile, NLP research on multilingual LLMs continues to prioritize high-resource languages. However, these models exhibit limited performance in under-resourced languages, primarily due to insufficient language- and culture-specific training data. In this paper, we focus on Modern Greek, for which only a limited number of question answering (QA) datasets have been proposed, most of which are intended for model evaluation. To address this research gap in Greek QA, we make the following contributions: (i) CulturaQA, a high-quality LRM-generated and human-curated dataset, for Greek LLM training and evaluation; (ii) a memory-efficient LLM evaluation framework adaptable to diverse languages and QA tasks; (iii) Maistros 8B, a state-of-the-art open-weights Greek LLM developed via knowledge distillation and fine-tuning on CulturaQA; and (iv) a comprehensive evaluation of nine LLMs across nine human-curated Greek QA datasets.

希腊语知识蒸馏小语种问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。