arXiv:2503.13102cs.CL2025-03被引 3

构建俄语错误标注数据集,评估大模型在俄语文本生成中的评判能力。

REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities

  • 构建1000个查询、2000条生成响应的俄语错误标注数据集REPA。
  • 发现俄语大模型评判者表现显著低于英语,但人类与模型偏好部分一致。
  • 适合关注多语言大模型评估、跨语言评测基准的研究者。

大语言模型作为评判者的新范式在英语中表现良好,但其在俄语中的应用尚未深入。本文引入俄罗斯错误类型标注数据集(REPA),包含1000个用户查询和2000条大模型生成响应。人工标注者对每对响应在十种具体错误类型上表达偏好,并选出整体偏好。我们基于人类偏好对六种生成模型按错误类型进行排名,并在零样本和少样本设置下评估八种大模型作为评判者的性能。分析显示,俄语大模型评判者表现明显弱于英语,但人类与模型偏好存在部分一致性,表明当前模型在俄语细粒度评估中仍有改进空间。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have introduced the novel paradigm of using LLMs as judges, where an LLM evaluates and scores the outputs of another LLM, which often correlates highly with human preferences. However, the use of LLM-as-a-judge has been primarily studied in English. In this paper, we evaluate this framework in Russian by introducing the Russian Error tyPes Annotation dataset (REPA), a dataset of 1k user queries and 2k LLM-generated responses. Human annotators labeled each response pair expressing their preferences across ten specific error types, as well as selecting an overall preference. We rank six generative LLMs across the error types using three rating systems based on human preferences. We also evaluate responses using eight LLM judges in zero-shot and few-shot settings. We describe the results of analyzing the judges and position and length biases. Our findings reveal a notable gap between LLM judge performance in Russian and English. However, rankings based on human and LLM preferences show partial alignment, suggesting that while current LLM judges struggle with fine-grained evaluation in Russian, there is potential for improvement.

大模型评估多语言生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。