构建菲律宾语评测基准,揭示大模型在菲语理解生成上的短板
FilBench: Can LLMs Understand and Generate Filipino?
- 设计涵盖文化知识、阅读理解等任务的菲律宾语评测集
- 最佳模型GPT-4o仅达72.23分,多数模型表现不佳
- 专为东南亚语言训练的模型在该基准上也表现平平
尽管大语言模型在英语任务中表现出色,但其在菲律宾语等特定语言上的能力仍不清楚。本文提出FilBench,一个以菲律宾语为核心的基准,用于评估大模型在菲律宾语、他加禄语和宿务语中的多项任务表现。任务设计覆盖文化知识、经典NLP、阅读理解与生成等菲律宾NLP研究重点。对27个主流大模型的评估显示,多个模型在阅读理解和翻译上表现较差。结果显示,FilBench具有挑战性,最优模型GPT-4o仅得72.23分;专门针对东南亚语言训练的模型(如SEA-LION v3 70B)表现更差,最高仅61.07分。本工作表明,构建语言专项基准对推动菲律宾语NLP发展及提升菲律宾语言在大模型中的包容性至关重要。
原文摘要 · Abstract (English)
Despite the impressive performance of LLMs on English-based tasks, little is known about their capabilities in specific languages such as Filipino. In this work, we address this gap by introducing FilBench, a Filipino-centric benchmark designed to evaluate LLMs across a diverse set of tasks and capabilities in Filipino, Tagalog, and Cebuano. We carefully curate the tasks in FilBench to reflect the priorities and trends of NLP research in the Philippines such as Cultural Knowledge, Classical NLP, Reading Comprehension, and Generation. By evaluating 27 state-of-the-art LLMs on FilBench, we find that several LLMs suffer from reading comprehension and translation capabilities. Our results indicate that FilBench is challenging, with the best model, GPT-4o, achieving only a score of 72.23%. Moreover, we also find that models trained specifically for Southeast Asian languages tend to underperform on FilBench, with the highest-performing model, SEA-LION v3 70B, achieving only a score of 61.07%. Our work demonstrates the value of curating language-specific LLM benchmarks to aid in driving progress on Filipino NLP and increasing the inclusion of Philippine languages in LLM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。