arXiv:2505.02177cs.CL2025-05被引 1

构建首个针对香港多任务语言理解的评测基准,评估中英粤三语能力。

Measuring Hong Kong Massive Multi-Task Language Understanding

  • 设计涵盖66个学科的2.67万道多选题,含粤语翻译任务
  • 最佳模型准确率仅75%,远低于主流评测表现
  • 适合研究多语种、跨文化AI模型的开发者与学者

多语言理解对大语言模型(LLM)的跨文化应用至关重要。然而,针对香港独特语言环境——以繁体中文书写、粤语为口语及文化背景——的评测基准仍不完善。为此,我们提出HKMMLU,一个评估香港语言能力和社科知识的多任务语言理解基准。该基准包含66个学科的26,698道多选题,分为四大类:科学、技术、工程与数学(STEM)、社会科学、人文学科及其他。此外,还纳入90,550个普通话-粤语翻译任务,以评估模型的多语言理解能力。我们在GPT-4o、Claude 3.7 Sonnet及18个不同规模的开源模型上进行了全面实验。结果显示,表现最佳的DeepSeek-V3模型准确率不足75%,显著低于MMLU和CMMLU水平。该差距凸显了现有模型在港式语言与知识领域能力的不足。我们进一步分析了问题语言、模型规模、提示策略、问题与推理令牌长度对性能的影响。预期HKMMLU将显著推动多语言与跨文化场景下大模型的发展,实现更广泛、更具影响力的落地应用。

原文摘要 · Abstract (English)

Multilingual understanding is crucial for the cross-cultural applicability of Large Language Models (LLMs). However, evaluation benchmarks designed for Hong Kong's unique linguistic landscape, which combines Traditional Chinese script with Cantonese as the spoken form and its cultural context, remain underdeveloped. To address this gap, we introduce HKMMLU, a multi-task language understanding benchmark that evaluates Hong Kong's linguistic competence and socio-cultural knowledge. The HKMMLU includes 26,698 multi-choice questions across 66 subjects, organized into four categories: Science, Technology, Engineering, and Mathematics (STEM), Social Sciences, Humanities, and Other. To evaluate the multilingual understanding ability of LLMs, 90,550 Mandarin-Cantonese translation tasks were additionally included. We conduct comprehensive experiments on GPT-4o, Claude 3.7 Sonnet, and 18 open-source LLMs of varying sizes on HKMMLU. The results show that the best-performing model, DeepSeek-V3, struggles to achieve an accuracy of 75\%, significantly lower than that of MMLU and CMMLU. This performance gap highlights the need to improve LLMs' capabilities in Hong Kong-specific language and knowledge domains. Furthermore, we investigate how question language, model size, prompting strategies, and question and reasoning token lengths affect model performance. We anticipate that HKMMLU will significantly advance the development of LLMs in multilingual and cross-cultural contexts, thereby enabling broader and more impactful applications.

多语言理解评测基准粤语模型跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。