评测大模型在多语种和混用语境下处理惯用语的能力,发现顶尖模型表现不佳。
Evaluating Large Language Models on Multiword Expressions in Multilingual and Code-Switched Contexts
- 构建跨语言与混用语数据集,测试模型对惯用语的识别与语义理解。
- GPT-4等最新模型在新任务上表现不如xlm-roBERTa-base基线模型。
- 惯用语歧义问题仍是大模型的难点,尤其在低频语境中表现更差。
惯用语具有非组合性语义和句法不规则性,是语言细微之处的典型代表。它们可字面或习语使用,导致语义显著变化。尽管大语言模型在诸多任务中表现出色,但其处理此类语言细微差别的能力仍不明确。本研究评估了前沿语言模型在较少见语境中处理潜在习语性惯用语歧义的能力,这些语境下模型难以依赖记忆。通过在葡萄牙语、加利西亚语和英语中进行评估,并使用新构建的混用语数据集与新任务,我们发现,尽管具备强大能力,大模型在处理细腻语言时仍面临挑战。尤其是,最新的模型(包括GPT-4)在检测与语义任务中均未超越xlm-roBERTa-base基线模型,且在新引入的任务上表现尤其差,尽管该任务与现有任务相似。总体而言,我们的结果表明,惯用语(特别是歧义性惯用语)依然是模型的难题。
原文摘要 · Abstract (English)
Multiword expressions, characterised by non-compositional meanings and syntactic irregularities, are an example of nuanced language. These expressions can be used literally or idiomatically, leading to significant changes in meaning. While large language models have demonstrated strong performance across many tasks, their ability to handle such linguistic subtleties remains uncertain. Therefore, this study evaluates how state-of-the-art language models process the ambiguity of potentially idiomatic multiword expressions, particularly in contexts that are less frequent, where models are less likely to rely on memorisation. By evaluating models across in Portuguese and Galician, in addition to English, and using a novel code-switched dataset and a novel task, we find that large language models, despite their strengths, struggle with nuanced language. In particular, we find that the latest models, including GPT-4, fail to outperform the xlm-roBERTa-base baselines in both detection and semantic tasks, with especially poor performance on the novel tasks we introduce, despite its similarity to existing tasks. Overall, our results demonstrate that multiword expressions, especially those which are ambiguous, continue to be a challenge to models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。