arXiv:2412.01186cs.CL2024-12被引 4

构建面向东南亚语言的可复现评估基准,助力大模型本地化性能评测

SailCompass: Towards Reproducible and Robust Evaluation for Southeast Asian Languages

  • 基于三种主要东南亚语言设计多任务评测体系
  • 专用模型仍优于通用模型,但差距缩小
  • 适配提示优化与校准技术,提升评测可靠性

本文提出SailCompass,一个针对东南亚语言(SEA)的大语言模型(LLM)可复现、鲁棒性评估基准。该基准涵盖三种主要东南亚语言,包含八项核心任务及14个数据集,覆盖生成、多项选择题和分类三类任务。为提升评估稳健性,研究探索了多项选择题的不同提示配置,并采用校准方法提升分类任务的准确性。实验发现:(1) 专用于东南亚语言的LLM仍优于通用模型,但差距已缩小;(2) 语言分布均衡对开发更优的SEA专用模型至关重要;(3) 高级提示技术(如校准、基于困惑度排序)能更好发挥模型潜力。所有数据集与评估脚本均已开源。

原文摘要 · Abstract (English)

In this paper, we introduce SailCompass, a reproducible and robust evaluation benchmark for assessing Large Language Models (LLMs) on Southeast Asian Languages (SEA). SailCompass encompasses three main SEA languages, eight primary tasks including 14 datasets covering three task types (generation, multiple-choice questions, and classification). To improve the robustness of the evaluation approach, we explore different prompt configurations for multiple-choice questions and leverage calibrations to improve the faithfulness of classification tasks. With SailCompass, we derive the following findings: (1) SEA-specialized LLMs still outperform general LLMs, although the gap has narrowed; (2) A balanced language distribution is important for developing better SEA-specialized LLMs; (3) Advanced prompting techniques (e.g., calibration, perplexity-based ranking) are necessary to better utilize LLMs. All datasets and evaluation scripts are public.

语言评估大模型东南亚语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。