arXiv:2605.27015cs.CL2026-05

评测大模型对波斯文学的细粒度理解能力,发现其在语法分析上表现差。

PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions

论文配图:PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
图 1 · 摘自论文原文
  • 构建4514道波斯文学多选题,覆盖8类细粒度知识
  • 模型在概念相似性任务准确率高,拼写与构词最差
  • 提示策略显著影响结果,解释型少样本效果最佳

尽管大语言模型具备出色的多语言能力,但在非英语文学知识上的评估仍不充分。我们提出PersLitEval,一个包含4,514道波斯文学多项选择题的基准,覆盖拼写、修辞手法、语法、词汇、构词和概念理解等八类细粒度类别,数据源自伊朗大学入学考试(Konkur)材料。我们在十种提示策略下评估六种LLMs,发现模型在概念相似性任务中表现较好,但在形式语言分析上表现不佳,拼写与构词问题最难。提示策略影响显著,解释型少样本提示在形式语言类别上效果最佳。错误分析识别出三类失败模式:语义理解缺失、形式语言知识不足及计数错误,表明不同类别需差异化改进。

原文摘要 · Abstract (English)

Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We introduce PersLitEval, a benchmark of 4,514 Persian literature multiple-choice questions across eight fine-grained categories spanning spelling, literary devices, grammar, vocabulary, word formation, and conceptual understanding, sourced from materials for the Konkur university entrance examination. We evaluate six LLMs across ten prompting strategies, revealing striking category-level disparities across three tiers of task difficulty: models reach higher accuracy on conceptual similarity tasks but struggle with formal linguistic analysis, with spelling and word formation proving the hardest across all models. Prompting strategy has a significant impact on performance, with explained few-shot examples yielding the best results, particularly on formal linguistic categories. An error analysis identifies three failure modes: semantic comprehension gaps, formal linguistic knowledge gaps, and counting/enumeration errors, suggesting that different categories require different improvement strategies.

文学理解波斯语细粒度评测大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。