用类比提示激发大模型在低资源语言中的语言推理能力
Inductive Linguistic Reasoning with Large Language Models
- 通过自动生成类比例证,增强模型对语言规律的归纳能力
- GPT-4o性能提升8.1%,Llama-3提升5.9%,超越传统思维链方法
- 适用于语言奥赛类任务,适合研究推理机制与多语言建模者
评估大语言模型(LLMs)在语言推理方面的能力,是理解其大规模应用中潜在短板的重要任务。本文聚焦于通过极低资源语言的语法谜题,考察模型进行抽象多语言推理的能力。由于翻译任务涉及从参考实例中进行归纳与演绎推理,我们探究是否可通过类比提示,从初始例证自动诱导出多样化的辅助示范。采用两阶段流程:首先由语言模型生成类比例证,再将其与目标语言例证一同作为上下文使用。在modeLing数据集上的实验表明,类比提示能有效激发模型对语言语法相似性的知识,使GPT-4o性能提升8.1%,Llama-3.1-405B-Instruct提升5.9%,优于链式思维方法。该增益源于自动生成或弱多语言模型生成的类比示范。此外,该方法在LINGOLY数据集上对语言奥赛各类问题均取得显著提升,覆盖不同难度层级。我们还发现了影响语言推理表现的若干有趣现象,表明此类谜题是检验新推理方法的宝贵基准。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) on their linguistic reasoning capabilities is an important task to understand the gaps in their skills that may surface during large-scale adoption. In this work, we investigate the abilities of such models to perform abstract multilingual reasoning through the lens of linguistic puzzles on extremely low-resource languages. As these translation tasks involve inductive and deductive reasoning from reference instances, we examine whether diverse auxiliary demonstrations can be automatically induced from seed exemplars, through analogical prompting. We employ a two-stage procedure, first generating analogical exemplars with a language model, and then applying them in-context along with provided target language exemplars. Our results on the modeLing dataset show that analogical prompting is effective in eliciting models' knowledge of language grammar similarities, boosting the performance of GPT-4o by as much as 8.1% and Llama-3.1-405B-Instruct by 5.9% over chain-of-thought approaches. These gains are attributable to the analogical demonstrations, both when self-generated as well as when produced by weaker multilingual models. Furthermore, we demonstrate that our method generalizes to other tasks present in Linguistics Olympiad competitions, achieving sizable improvements across all problem types and difficulty levels included in the LINGOLY dataset with GPT-4o. We also report several findings about interesting phenomena which drive linguistic reasoning performance, suggesting that such puzzles are a valuable benchmark for new reasoning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。