arXiv:2409.08564cs.CL2024-09NAACL被引 6

评测大模型在印尼真实职业考试中的表现,发现本地化领域效果差。

Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia

  • 构建印尼多领域职业考试数据集IndoCareer,含8834道题。
  • 大模型在保险金融等本地化强的领域表现显著不佳。
  • 答案顺序打乱影响保险金融类评估结果稳定性,适合本地化研究者参考。

尽管大语言模型的知识评估主要集中在数学、物理等学术科目,但这些测试往往无法反映真实职业的实际需求。本文提出IndoCareer数据集,包含8,834道多项选择题,用于评估大模型在印尼多个职业认证考试中的表现。该数据集覆盖六个关键领域:(1)医疗健康,(2)保险与金融,(3)创意设计,(4)旅游与酒店业,(5)教育与培训,(6)法律。对27个大语言模型的全面评估显示,模型在依赖本地语境的领域(如保险与金融)表现较差。此外,在使用完整数据集时,打乱选项顺序通常能保持各模型评估结果一致,但在保险与金融领域却引发评估结果不稳定。

原文摘要 · Abstract (English)

While knowledge evaluation in large language models has predominantly focused on academic subjects like math and physics, these assessments often fail to capture the practical demands of real-world professions. In this paper, we introduce IndoCareer, a dataset comprising 8,834 multiple-choice questions designed to evaluate performance in vocational and professional certification exams across various fields. With a focus on Indonesia, IndoCareer provides rich local contexts, spanning six key sectors: (1) healthcare, (2) insurance and finance, (3) creative and design, (4) tourism and hospitality, (5) education and training, and (6) law. Our comprehensive evaluation of 27 large language models shows that these models struggle particularly in fields with strong local contexts, such as insurance and finance. Additionally, while using the entire dataset, shuffling answer options generally maintains consistent evaluation results across models, but it introduces instability specifically in the insurance and finance sectors.

职业评估本地化多领域大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。