arXiv:2607.28274cs.CLcs.LG2026-07被引 1

首个专测现代希腊语词形变化能力的基准,揭示大模型在形态丰富语言上的短板。

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

  • 构建500个专家验证题,测试希腊语词形识别与生成能力,侧重低频词以考规则理解。
  • 评估从LLaMA到Kimi K2等多模型,发现多数在词形变化上表现不佳,即使多语言覆盖强。
  • 自研模型Sophea-Genesis-1在词形任务领先,且通用能力与同规模模型相当。

现代希腊语是一种高度屈折的语言,但现有针对该语言的模型主要在事实知识上进行评估,缺乏专门衡量其词形变化能力的基准。本文提出MORFES(形态开放类识别与生成评估套件),包含500个由专家验证的题目,用于测试对希腊语屈折形式的识别与生成能力,并优先选择低频词项,确保正确答案反映规则而非记忆。该数据集已公开于https://huggingface.co/datasets/KIEFERSA/MORFES。我们在多个开源语言模型(涵盖从LLaMA到Qwen3、DeepSeek-R1、Magistral及Kimi K2)上评估其表现,这些模型虽具备日益增强的多语言能力,但在形态丰富的语言中语法能力仍被低估。其中,我们自研并开源的Sophea-Genesis-1模型在词形变化任务中表现最优,同时在通用能力上与同规模模型相当。

原文摘要 · Abstract (English)

Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicated to their inflectional competence. We introduce MORFES (Morphological Open-class Recognition-and-Formation Evaluation Suite), a benchmark of 500 expert-verified items that tests the recognition and production of Greek inflected forms, favoring lower-frequency lemmas so that a correct answer reflects the rule rather than a memorized form. We make it publicly available at https://huggingface.co/datasets/KIEFERSA/MORFES. We evaluate a range of open language models on MORFES, situating them within the rapidly scaling open-weight ecosystem from LLaMA to Qwen3, DeepSeek-R1, Magistral, and Kimi K2, where multilingual coverage grows but grammatical competence in morphologically rich languages remains under-measured. Among them, Sophea-Genesis-1, a model we developed and release as open weights at https://huggingface.co/KIEFERSA/Sophea-Genesis-1, leads on inflectional morphology while matching similarly sized models in general capability.

语言模型形态学希腊语评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。