arXiv:2602.22827cs.CLcs.LG2026-02被引 1

构建首个波斯语文化理解评估基准,提升大模型评测准确性。

TARAZ: Persian Short-Answer Question Benchmark for Cultural Evaluation of Language Models

  • 设计波斯语短答案评估框架,融合形态归一化与混合语义相似度计算。
  • 相比精确匹配基线,评分一致性提升10分,更准确捕捉语义差异。
  • 适合研究跨文化语言模型能力或波斯语NLP的学者使用。

本文提出一个全面的评估框架,用于衡量大语言模型(LLMs)在波斯语中的文化胜任力。现有波斯语文化基准多依赖多项选择题形式和以英语为中心的评价指标,难以反映波斯语的形态复杂性和语义细微差别。我们的框架引入专门针对波斯语的短答案评估,结合基于规则的形态归一化与混合句法-语义相似度模块,实现超越精确字符串匹配的稳健软匹配评分。通过在三个文化相关的波斯语数据集上系统评估15个最先进的开源与闭源模型,我们证明该混合评估方法在评分一致性上比精确匹配基线提高+10分,能有效捕捉表面方法无法识别的意义。人工评估进一步验证,所提出的语义相似度度量与人类判断的一致性高于基于LLM的评判者。我们公开发布该评估框架,提供首个标准化基准以衡量波斯语中的文化理解能力,并为跨文化大模型评估研究建立可复现的基础。

原文摘要 · Abstract (English)

This paper presents a comprehensive evaluation framework for assessing the cultural competence of large language models (LLMs) in Persian. Existing Persian cultural benchmarks rely predominantly on multiple-choice formats and English-centric metrics that fail to capture Persian's morphological complexity and semantic nuance. Our framework introduces a Persian-specific short-answer evaluation that combines rule-based morphological normalization with a hybrid syntactic and semantic similarity module, enabling robust soft-match scoring beyond exact string overlap. Through systematic evaluation of 15 state-of-the-art open- and closed-source models across three culturally grounded Persian datasets, we demonstrate that our hybrid evaluation improves scoring consistency by +10 compared to exact-match baselines by capturing meaning that surface-level methods cannot detect. Our human evaluation further confirms that the proposed semantic similarity metric achieves higher agreement with human judgments than LLM-based judges. We publicly release our evaluation framework, providing the first standardized benchmark for measuring cultural understanding in Persian and establishing a reproducible foundation for cross-cultural LLM evaluation research.

文化理解波斯语评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。