为欧洲语言大模型评估建立新分类体系与实践标准
Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
- 提出专用于多语言场景的评测基准分类法
- 强调评测需提升对语言文化敏感度
- 适合关注欧洲语言AI评估的研究者
尽管大型语言模型(LLMs)的新评测基准持续涌现以应对模型能力的快速演进,但非英语语言的LLM评估仍处于探索阶段。本文简要综述了近期LLM评测的发展,提出一种针对多语言或非英语使用场景的新型评测基准分类体系,并进一步提出一套最佳实践与质量标准,旨在推动欧洲语言评测基准的协同建设。建议包括提升评估方法的语言与文化敏感性,确保评测结果更具代表性与公平性。
原文摘要 · Abstract (English)
While new benchmarks for large language models (LLMs) are being developed continuously to catch up with the growing capabilities of new models and AI in general, using and evaluating LLMs in non-English languages remains a little-charted landscape. We give a concise overview of recent developments in LLM benchmarking, and then propose a new taxonomy for the categorization of benchmarks that is tailored to multilingual or non-English use scenarios. We further propose a set of best practices and quality standards that could lead to a more coordinated development of benchmarks for European languages. Among other recommendations, we advocate for a higher language and culture sensitivity of evaluation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。