arXiv:2410.11216cs.CL2024-10

构建英式、澳式、印式英语情感分类基准,揭示语言差异对模型性能的影响。

Experiences from Creating a Benchmark for Sentiment Classification for Varieties of English

  • 基于谷歌地点评论,设计多语言变体采样策略
  • 不同语言变体下模型性能差异显著,最高达23%波动
  • 适合关注跨语言NLP与公平性评估的研究者

现有基准常忽略英语的语言多样性。本文分享了构建英式(en-AU)、澳式(en-AU)和印式(en-IN)英语情感分类基准的实践经验。基于Google Places评论,我们探索了基于标签语义、评论长度和情感比例的多种采样方法,并在三个微调的BERT模型上报告结果。初步评估显示,样本特征、标签语义和语言变体显著影响性能,最大性能波动达23%。研究强调需采用多样化采样、谨慎定义标签,并在多语言变体中进行全面评估,为构建稳健基准提供可操作建议。

原文摘要 · Abstract (English)

Existing benchmarks often fail to account for linguistic diversity, like language variants of English. In this paper, we share our experiences from our ongoing project of building a sentiment classification benchmark for three variants of English: Australian (en-AU), Indian (en-IN), and British (en-UK) English. Using Google Places reviews, we explore the effects of various sampling techniques based on label semantics, review length, and sentiment proportion and report performances on three fine-tuned BERT-based models. Our initial evaluation reveals significant performance variations influenced by sample characteristics, label semantics, and language variety, highlighting the need for nuanced benchmark design. We offer actionable insights for researchers to create robust benchmarks, emphasising the importance of diverse sampling, careful label definition, and comprehensive evaluation across linguistic varieties.

情感分析语言变体数据采样BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。