arXiv:2502.07058cs.CLcs.HC2025-02NAACL

用在线评论对比中英文变体下大模型表现差异

Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties

  • 利用同一酒店的两岸中文评论构建对齐数据集
  • 六款大模型在台湾中文上表现均显著下降
  • 为评估语言多样性影响提供低成本新方法

一种语言可存在多种变体,这些变体会影响自然语言处理模型(包括大语言模型)的性能,而大模型通常在主流变体数据上训练。本文提出一种新颖且成本低廉的方法,用于跨语言变体评估模型性能。我们论证国际在线评论平台(如Booking.com)可作为有效数据源,构建包含不同语言变体的评论数据集,这些评论来自相似真实场景(如同一酒店、相同评分、同种语言,但使用不同变体,如台湾中文与大陆中文)。为验证该方法,我们构建了包含台湾中文与大陆中文评论的上下文对齐数据集,并在情感分析任务中测试了六款大语言模型。结果表明,大模型在台湾中文上的表现持续低于大陆中文。

原文摘要 · Abstract (English)

A language can have different varieties. These varieties can affect the performance of natural language processing (NLP) models, including large language models (LLMs), which are often trained on data from widely spoken varieties. This paper introduces a novel and cost-effective approach to benchmark model performance across language varieties. We argue that international online review platforms, such as Booking.com, can serve as effective data sources for constructing datasets that capture comments in different language varieties from similar real-world scenarios, like reviews for the same hotel with the same rating using the same language (e.g., Mandarin Chinese) but different language varieties (e.g., Taiwan Mandarin, Mainland Mandarin). To prove this concept, we constructed a contextually aligned dataset comprising reviews in Taiwan Mandarin and Mainland Mandarin and tested six LLMs in a sentiment analysis task. Our results show that LLMs consistently underperform in Taiwan Mandarin.

大模型评测语言变体情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。