arXiv:2602.18788cs.CL2026-02被引 1

首个系统评估缅甸语大模型的综合基准,覆盖理解、推理与生成三方面。

BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models

  • 基于母语者构建,确保语言自然真实,减少翻译偏差。
  • 结果显示模型架构和指令微调比规模更影响缅甸语表现。
  • 适合关注低资源语言、东南亚语言模型的研究者使用。

我们提出了 BURMESE-SAN,首个系统评估大语言模型在缅甸语上表现的综合性基准,涵盖自然语言理解(NLU)、自然语言推理(NLR)和自然语言生成(NLG)三大核心能力。该基准整合了七个子任务:问答、情感分析、毒性检测、因果推理、自然语言推断、抽象摘要和机器翻译,其中多项任务此前在缅甸语中尚未可用。基准通过严格的母语者驱动流程构建,确保语言自然性、流畅性和文化真实性,同时最大限度减少翻译引入的偏差。我们对开源与商用大模型进行了大规模评估,揭示了缅甸语建模面临的挑战:预训练覆盖不足、形态丰富及句法多变。结果表明,缅甸语性能更多取决于模型架构设计、语言表征能力和指令微调,而非单纯模型规模。特别是东南亚区域微调和新世代模型带来显著提升。最后,我们公开发布 BURMESE-SAN 公共排行榜,以支持缅甸语及其他低资源语言的持续系统评估与进步。https://leaderboard.sea-lion.ai/detailed/MY

原文摘要 · Abstract (English)

We introduce BURMESE-SAN, the first holistic benchmark that systematically evaluates large language models (LLMs) for Burmese across three core NLP competencies: understanding (NLU), reasoning (NLR), and generation (NLG). BURMESE-SAN consolidates seven subtasks spanning these competencies, including Question Answering, Sentiment Analysis, Toxicity Detection, Causal Reasoning, Natural Language Inference, Abstractive Summarization, and Machine Translation, several of which were previously unavailable for Burmese. The benchmark is constructed through a rigorous native-speaker-driven process to ensure linguistic naturalness, fluency, and cultural authenticity while minimizing translation-induced artifacts. We conduct a large-scale evaluation of both open-weight and commercial LLMs to examine challenges in Burmese modeling arising from limited pretraining coverage, rich morphology, and syntactic variation. Our results show that Burmese performance depends more on architectural design, language representation, and instruction tuning than on model scale alone. In particular, Southeast Asia regional fine-tuning and newer model generations yield substantial gains. Finally, we release BURMESE-SAN as a public leaderboard to support systematic evaluation and sustained progress in Burmese and other low-resource languages. https://leaderboard.sea-lion.ai/detailed/MY

缅甸语大模型评估低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。