arXiv:2602.15373cs.CLcs.AI2026-02中稿 · 13th VarDial works…

评测大模型对印式与澳式英语俚语的理解能力,发现性能差异显著。

Far Out: Evaluating Language Models on Slang in Australian and Indian English

  • 构建网页和生成两类数据集,覆盖多种俚语使用场景。
  • 识别任务准确率最高(0.54),生成类任务表现最差(0.03)。
  • 印式英语俚语理解优于澳式英语,尤其在选择任务中差距明显。

语言模型在处理非标准语料时存在系统性性能差距,但对特定方言俚语的理解能力仍缺乏深入研究。本文针对印度英语(en-IN)和澳大利亚英语(en-AU)开展全面评估,涵盖七种先进语言模型。构建两个互补数据集:WEB(来自Urban Dictionary的377条网络来源俚语实例)和GEN(1,492条合成生成的俚语用例),覆盖多样情境。评估三种任务:目标词预测(TWP)、引导式目标词预测(TWP$^*$)和目标词选择(TWS)。结果显示:(1)TWS平均准确率最高(0.49),远高于TWP(0.03)和TWP$^*$(0.03);(2)WEB数据集上模型表现优于GEN,TWP与TWP$^*$任务相似度分别提升0.03和0.05;(3)en-IN任务整体优于en-AU,TWS任务中平均准确率从0.44升至0.54。结果揭示生成与判别能力在方言俚语理解中的根本不对称性。

原文摘要 · Abstract (English)

Language models exhibit systematic performance gaps when processing text in non-standard language varieties, yet their ability to comprehend variety-specific slang remains underexplored for several languages. We present a comprehensive evaluation of slang awareness in Indian English (en-IN) and Australian English (en-AU) across seven state-of-the-art language models. We construct two complementary datasets: WEB, containing 377 web-sourced usage examples from Urban Dictionary, and GEN, featuring 1,492 synthetically generated usages of these slang terms, across diverse scenarios. We assess language models on three tasks: target word prediction (TWP), guided target word prediction (TWP$^*$) and target word selection (TWS). Our results reveal four key findings: (1) Higher average model performance TWS versus TWP and TWP$^*$, with average accuracy score increasing from 0.03 to 0.49 respectively (2) Stronger average model performance on WEB versus GEN datasets, with average similarity score increasing by 0.03 and 0.05 across TWP and TWP$^*$ tasks respectively (3) en-IN tasks outperform en-AU when averaged across all models and datasets, with TWS demonstrating the largest disparity, increasing average accuracy from 0.44 to 0.54. These findings underscore fundamental asymmetries between generative and discriminative competencies for variety-specific language, particularly in the context of slang expressions despite being in a technologically rich language such as English.

语言模型俚语理解方言差异多语言评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。