arXiv:2411.01706cs.CL2024-11EMNLP被引 11

大模型在多语言复杂词识别任务中表现一般,难超小型专用模型。

Investigating Large Language Models for Complex Word Identification in Multilingual and Multidomain Setups

  • 对比Llama、Vicuna、ChatGPT等10余款大小模型在多场景下的表现
  • 零样本与少样本设置下准确率普遍低于传统小模型,最高仅达76.2%
  • 适合研究多语言文本简化或对大模型泛化能力感兴趣的读者

复杂词识别(CWI)是词汇简化中的关键步骤,近年已发展为独立任务。该任务有多种变体,如词汇复杂度预测(LCP)和多词表达式复杂度评估(MWE)。大语言模型(LLMs)因具备零/少样本泛化能力,在自然语言处理领域备受关注。本文研究了开源模型(Llama 2、Llama 3、Vicuna v1.5)与闭源模型(ChatGPT-3.5-turbo、GPT-4o)在CWI、LCP和MWE任务中的表现,涵盖零样本、少样本及微调设置。实验表明,尽管部分模型在特定条件下表现尚可,但整体上难以超越现有小型专用模型,尤其在跨语言、跨领域场景中性能下降明显。同时,本文探讨了元学习结合提示学习的潜力。结论指出,当前主流大模型尚未能在复杂词识别任务中显著优于轻量级方法。

原文摘要 · Abstract (English)

Complex Word Identification (CWI) is an essential step in the lexical simplification task and has recently become a task on its own. Some variations of this binary classification task have emerged, such as lexical complexity prediction (LCP) and complexity evaluation of multi-word expressions (MWE). Large language models (LLMs) recently became popular in the Natural Language Processing community because of their versatility and capability to solve unseen tasks in zero/few-shot settings. Our work investigates LLM usage, specifically open-source models such as Llama 2, Llama 3, and Vicuna v1.5, and closed-source, such as ChatGPT-3.5-turbo and GPT-4o, in the CWI, LCP, and MWE settings. We evaluate zero-shot, few-shot, and fine-tuning settings and show that LLMs struggle in certain conditions or achieve comparable results against existing methods. In addition, we provide some views on meta-learning combined with prompt learning. In the end, we conclude that the current state of LLMs cannot or barely outperform existing methods, which are usually much smaller.

复杂词识别大模型评测多语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。