构建首个针对英式方言的情感与反讽分类基准,助力公平化大模型研究。
BESSTIE: A Benchmark for Sentiment and Sarcasm Classification for Varieties of English
- 采集澳、印、英三种英语变体数据,结合地理位置与主题筛选。
- 模型对印度英语的反讽识别准确率显著低于英澳英语。
- 适合关注语言多样性与模型公平性的研究人员使用。
尽管大语言模型(LLMs)被发现对非标准英语变体存在偏见,但目前缺乏针对英语的情感分析标注数据集。为填补这一空白,我们提出BESSTIE,一个面向澳大利亚(en-AU)、印度(en-IN)和英国(en-UK)三种英语变体的情感与反讽分类基准。通过基于地理位置的Google Places评论和基于主题的Reddit评论筛选方法收集数据,并进行两轮验证:(a)母语者手动标注语言变体;(b)自动语言变体预测。母语者对数据集进行情感与反讽标注,并开展额外标注验证。随后,我们在九个不同架构的LLM(涵盖编码器/解码器及单语/多语模型)上微调并评估其性能。结果显示,模型在内圈英语变体(en-AU与en-UK)上的表现优于en-IN,尤其在反讽分类任务中差异明显。同时揭示跨变体泛化困难,凸显建立语言变体专用数据集的必要性。BESSTIE可作为未来公平化大模型研究的重要评估基准。数据集已公开于:https://huggingface.co/datasets/unswnlporg/BESSTIE。
原文摘要 · Abstract (English)
Despite large language models (LLMs) being known to exhibit bias against non-standard language varieties, there are no known labelled datasets for sentiment analysis of English. To address this gap, we introduce BESSTIE, a benchmark for sentiment and sarcasm classification for three varieties of English: Australian (en-AU), Indian (en-IN), and British (en-UK). We collect datasets for these language varieties using two methods: location-based for Google Places reviews, and topic-based filtering for Reddit comments. To assess whether the dataset accurately represents these varieties, we conduct two validation steps: (a) manual annotation of language varieties and (b) automatic language variety prediction. Native speakers of the language varieties manually annotate the datasets with sentiment and sarcasm labels. We perform an additional annotation exercise to validate the reliance of the annotated labels. Subsequently, we fine-tune nine LLMs (representing a range of encoder/decoder and mono/multilingual models) on these datasets, and evaluate their performance on the two tasks. Our results show that the models consistently perform better on inner-circle varieties (i.e., en-AU and en-UK), in comparison with en-IN, particularly for sarcasm classification. We also report challenges in cross-variety generalisation, highlighting the need for language variety-specific datasets such as ours. BESSTIE promises to be a useful evaluative benchmark for future research in equitable LLMs, specifically in terms of language varieties. The BESSTIE dataset is publicly available at: https://huggingface.co/ datasets/unswnlporg/BESSTIE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。