首个罗马尼亚语语法推理与解释基准,评测大模型教育能力
GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
- 构建1151道罗马尼亚高考题构成的多选题基准
- 谷歌Gemini 2.5 Pro准确率达83%,多数开源模型低于65%
- 揭示形态学与新拼写规范应用中的系统性错误
大语言模型(LLMs)已革新自然语言处理,但其在低资源语言中的教学价值尚不明确。我们提出GRILE(罗马尼亚语语法推理与解释基准),首个公开的1,151道多选题数据集,源自罗马尼亚高利害考试(国家评估、毕业考、大学入学)。该基准用于评估七种先进多语言及罗马尼亚专用大模型的两种能力:(i) 正确选项选择,(ii) 生成语言学上准确的解释。尽管Gemini 2.5 Pro达到83%准确率,多数开源模型仍低于65%,且48%的解释经专家评审存在事实或教学错误。详细错误分析揭示了形态学和最新DOOM3拼写规范应用中的系统性缺陷。所有数据、代码及公开网页演示均已发布,以推动未来研究。研究结果揭示了低资源教育NLP中的可信挑战,并确立GRILE作为可控解释生成与评估的新测试平台。
原文摘要 · Abstract (English)
LLMs (Large language models) have revolutionized NLP (Natural Language Processing), yet their pedagogical value for low-resource languages remains unclear. We present GRILE (Grammar Romanian Inference and Language Explanations) , the first open benchmark of 1,151 multiple-choice questions harvested from Romanian high-stakes exams (National Evaluation, Baccalaureate, university admissions). GRILE enables us to probe two complementary abilities of seven state-of-the-art multilingual and Romanian-specific LLMs: (i) selecting the correct answer, and (ii) producing linguistically accurate explanations. While Gemini 2.5 Pro reaches 83% accuracy, most open-weight models stay below 65%, and 48% of their explanations contain factual or pedagogical flaws according to expert review. A detailed error analysis pinpoints systematic weaknesses in morphology and in applying the latest DOOM3 orthographic norms. All data, code and a public web demo are released to catalyze future research. Our findings expose open challenges for trustworthy educational NLP in low-resource settings and establish GRILE as a new test-bed for controllable explanation generation and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。