对比多种提示策略,提升低资源语言数据生成效果
A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages
- 组合使用目标语言示例与大模型修订,生成效果更优
- 部分设置下合成数据性能逼近真实数据仅差5%
- 小模型也能通过智能提示实现高效数据生成
大型语言模型(LLMs)越来越多地用于生成合成文本数据以训练小型专用模型。然而,针对低资源语言环境下的各类生成策略对比仍不充分。尽管已提出多种提示策略,如示范、基于标签的摘要和自修订,但它们在低资源语言中的相对有效性尚不明确。本文系统评估了这些生成策略及其组合在11种语言上的表现,涵盖多种类型差异显著的语言,包括若干极度低资源语言。在三个NLP任务和四个开源LLM上,评估下游模型在生成数据与真实数据上的性能。结果表明,战略性组合生成方法,特别是目标语言示范结合基于LLM的修订,能取得优异效果,使合成数据与真实数据的性能差距缩小至部分场景下的5%。同时发现,智能提示技术可降低大模型的优势,表明在低资源场景中,小模型也可通过高效生成策略实现良好表现。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to generate synthetic textual data for training smaller specialized models. However, a comparison of various generation strategies for low-resource language settings is lacking. While various prompting strategies have been proposed, such as demonstrations, label-based summaries, and self-revision, their comparative effectiveness remains unclear, especially for low-resource languages. In this paper, we systematically evaluate the performance of these generation strategies and their combinations across 11 typologically diverse languages, including several extremely low-resource ones. Using three NLP tasks and four open-source LLMs, we assess downstream model performance on generated versus gold-standard data. Our results show that strategic combinations of generation methods, particularly target-language demonstrations with LLM-based revisions, yield strong performance, narrowing the gap with real data to as little as 5% in some settings. We also find that smart prompting techniques can reduce the advantage of larger LLMs, highlighting efficient generation strategies for synthetic data generation in low-resource scenarios with smaller models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。