arXiv:2503.12440cs.CL2025-03被引 10

评测大模型对粤语及香港文化的理解能力,发现现有模型仍存明显短板。

HKCanto-Eval: A Benchmark for Evaluating Cantonese Language Understanding and Cultural Comprehension in LLMs

  • 构建包含粤语、英语和中文的跨语言评估基准
  • 在粤语文化理解任务上,闭源模型表现优于开源模型
  • 特别设计问题考察模型对香港语言背景知识的掌握程度

大型语言模型在多元语言与文化环境中的理解与交互能力至关重要。香港使用的粤语因富含文化细节且缺乏专用评估数据集,给自然语言处理带来独特挑战。本文提出HKCanto-Eval基准,用于评估大语言模型在粤语理解任务上的表现,并扩展至英语与书面中文进行跨语言对比。该基准融合了香港特有的语言与文化细节,为模型在真实场景下的表现提供可靠评估框架。此外,基准中包含旨在探测模型潜在语言元知识的问题。实验结果表明,尽管闭源模型整体优于开源模型,但在处理粤语特有语言与文化知识方面仍存在显著不足,凸显出需要更针对性的训练数据与评估方法。代码已开源:https://github.com/hon9kon9ize/hkeval2025

原文摘要 · Abstract (English)

The ability of language models to comprehend and interact in diverse linguistic and cultural landscapes is crucial. The Cantonese language used in Hong Kong presents unique challenges for natural language processing due to its rich cultural nuances and lack of dedicated evaluation datasets. The HKCanto-Eval benchmark addresses this gap by evaluating the performance of large language models (LLMs) on Cantonese language understanding tasks, extending to English and Written Chinese for cross-lingual evaluation. HKCanto-Eval integrates cultural and linguistic nuances intrinsic to Hong Kong, providing a robust framework for assessing language models in realistic scenarios. Additionally, the benchmark includes questions designed to tap into the underlying linguistic metaknowledge of the models. Our findings indicate that while proprietary models generally outperform open-weight models, significant limitations remain in handling Cantonese-specific linguistic and cultural knowledge, highlighting the need for more targeted training data and evaluation methods. The code can be accessed at https://github.com/hon9kon9ize/hkeval2025

粤语理解文化认知评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。