arXiv:2603.16406cs.CLcs.AI2026-03中稿 · LREC 2026

揭露冰岛语大模型评测中的数据缺陷,呼吁低资源语言评估需人工验证。

Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic

  • 对比人工撰写与机器生成数据的评测效果差异
  • 发现未验证的合成/翻译数据存在严重错误,影响评测可信度
  • 适合关注低资源语言评测可靠性的研究者阅读

本文评估了当前冰岛语大语言模型的基准测试方法,揭示其存在的问题,并呼吁改进低/中资源语言的评估体系。研究发现,未经验证的合成或机器翻译数据常包含严重错误,可能扭曲评测结果并削弱测试有效性。定量分析表明,人工撰写或翻译的基准与合成/机器翻译基准之间存在显著差异。在当前机器翻译质量有限的情况下,依赖未验证的翻译数据会引入不可控误差,因此建议在低/中资源语言场景下必须进行人工验证,以确保评测可靠性。

原文摘要 · Abstract (English)

This paper evaluates current Large Language Model (LLM) benchmarking for Icelandic, identifies problems, and calls for improved evaluation methods in low/medium-resource languages in particular. We show that benchmarks that include synthetic or machine-translated data that have not been verified in any way, commonly contain severely flawed test examples that are likely to skew the results and undermine the tests' validity. We warn against the use of such methods without verification in low/medium-resource settings as the translation quality can, at best, only be as good as MT quality for a given language at any given time. Indeed, the results of our quantitative error analysis on existing benchmarks for Icelandic show clear differences between human-authored/-translated benchmarks vs. synthetic or machine-translated benchmarks.

大模型评测低资源语言冰岛语数据验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。