用提示工程从大模型中挖出西非语言的可用文本数据
Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe

- 设计六种提示策略,从商用大模型提取目标语言文本
- GPT-4o Mini 每次调用产出的可用词数是 Gemini 的6至41倍
- 不同语言需适配不同提示:豪萨语用对话,丰贝语用约束生成
大型语言模型(LLMs)训练数据来自低资源语言社区,但其内含的语言知识仅可通过商业API访问。本文研究是否可通过策略性提示从大模型中提取可用于豪萨语(非洲-亚细亚语系,约8000万使用者)和丰贝语(尼日尔-刚果语系,约200万使用者)的可用文本数据。我们在两个商用大模型(GPT-4o Mini 和 Gemini 2.5 Flash)上系统比较了六种提示任务类型。结果显示,GPT-4o Mini 每次API调用提取的可用目标语言词汇量比 Gemini 多6至41倍。最优策略因语言而异:豪萨语受益于功能性文本与对话提示,丰贝语则需使用约束生成提示。所有生成语料库与代码均已公开。
原文摘要 · Abstract (English)
Large language models (LLMs) are trained on data contributed by low-resource language communities, yet the linguistic knowledge encoded in these models remains accessible only through commercial APIs. This paper investigates whether strategic prompting can extract usable text data from LLMs for two West African languages: Hausa (Afroasiatic, approximately 80 million speakers) and Fongbe (Niger-Congo, approximately 2 million speakers). We systematically compare six elicitation task types across two commercial LLMs (GPT-4o Mini and Gemini 2.5 Flash). GPT-4o Mini extracts 6-41 times more usable target-language words per API call than Gemini. Optimal strategies differ by language: Hausa benefits from functional text and dialogue, while Fongbe requires constrained generation prompts. We release all generated corpora and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。