arXiv:2603.09979cs.CL2026-03被引 1

测试大模型对波斯诗和莎士比亚十四行诗的精确诗句召回能力

GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets

  • 构建波斯柔巴依诗与莎士比亚十四行诗的语义/词汇补全与识别基准
  • 模型诗句补全率低,但识别能力较强,诗人差异影响表现
  • 揭示后训练可能削弱对经典文本的直接记忆,适合文化计算研究者

波斯诗歌在伊朗文化中广泛使用,哈菲兹、萨迪等名家诗句常被引述、改写或根据不完整线索补全。这要求语言模型能基于语义和词汇信息准确调用经典诗句原文。我们提出GhazalBench,一个评估大语言模型(LLMs)在真实使用场景下获取波斯柔巴依诗经典句式的基准。不同于以往将记忆视为缺陷,本工作聚焦于精确文本调用的实际价值。基准涵盖多种语义与词汇提示下的补全与识别任务。在多款商用与开源多语言LLM上,精确补全仍具挑战性,而识别性能显著更高。不同诗人与模型家族间表现差异明显。对莎士比亚十四行诗的平行实验显示部分模型补全率更高,与训练数据中经典文本暴露程度有关。附加实验表明后训练可能降低模型对经典文本的直接续写能力。研究呼吁区分生成与识别的评估框架,并重视对文化重要文本的访问能力。GhazalBench已开源:https://github.com/kalhorghazal/GhazalBench。

原文摘要 · Abstract (English)

Persian poetry plays an active role in Iranian cultural practice, where verses by canonical poets such as Hafez and Saadi are frequently quoted, paraphrased, or completed from incomplete cues. Supporting such interactions requires language models to reliably access canonical verses from semantic and lexical information. We introduce GhazalBench, a benchmark for evaluating how large language models (LLMs) access the canonical surface forms of Persian ghazal under usage-grounded conditions. Unlike prior work that primarily treats memorization as a liability, GhazalBench studies settings in which access to exact wording is functionally important. The benchmark evaluates completion and recognition under varied semantic and lexical cues. Across proprietary and open-weight multilingual LLMs, exact completion remains challenging, while recognition is substantially stronger. Performance also varies across poets and model families. Parallel experiments on Shakespearean sonnets yield markedly higher completion for several models, consistent with differences in exposure to canonical texts. {Additional experiments suggest that post-training may reduce direct continuation-based access to canonical text.} Our findings motivate evaluation frameworks that distinguish production from recognition and assess access to culturally significant canonical texts. GhazalBench is available at https://github.com/kalhorghazal/GhazalBench.

诗歌生成文化计算语言模型评测波斯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。