arXiv:2409.07054cs.CLcs.AI2024-09被引 15

对比母语与非母语提示对大模型表现的影响,发现非母语提示更优。

Native vs Non-Native Language Prompting: A Comparative Analysis

  • 在12个阿拉伯语数据集上测试三种提示策略:母语、非母语、混合
  • 平均来看,非母语提示效果最佳,优于母语和混合提示
  • 适用于低资源语言场景下的提示工程研究

大型语言模型(LLMs)在自然语言处理(NLP)任务中表现出色。为激发模型知识,提示(prompt)起关键作用,通常由自然语言指令构成。大多数开源与闭源的LLM基于文本、图像、音频和视频等数字内容进行训练,因此在高资源语言上表现更好,而在中低资源语言上表现较差。由于提示在理解模型能力方面至关重要,提示所用语言成为重要研究问题。尽管已有相关研究,但针对中低资源语言的探索仍有限。本研究在11个不同NLP任务和12个阿拉伯语数据集(共9.7K数据点)上,对3种提示策略(母语、非母语、混合)进行了系统评估。共执行197次实验,涵盖3个LLM。结果表明,平均而言,非母语提示表现最佳,其次为混合提示,母语提示最差。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown remarkable abilities in different fields, including standard Natural Language Processing (NLP) tasks. To elicit knowledge from LLMs, prompts play a key role, consisting of natural language instructions. Most open and closed source LLMs are trained on available labeled and unlabeled resources--digital content such as text, images, audio, and videos. Hence, these models have better knowledge for high-resourced languages but struggle with low-resourced languages. Since prompts play a crucial role in understanding their capabilities, the language used for prompts remains an important research question. Although there has been significant research in this area, it is still limited, and less has been explored for medium to low-resourced languages. In this study, we investigate different prompting strategies (native vs. non-native) on 11 different NLP tasks associated with 12 different Arabic datasets (9.7K data points). In total, we conducted 197 experiments involving 3 LLMs, 12 datasets, and 3 prompting strategies. Our findings suggest that, on average, the non-native prompt performs the best, followed by mixed and native prompts.

提示工程低资源语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。