用合成数据让小模型学会印地语旅游问答,解决低资源语言难题
Adapting Small Language Models to Low-Resource Domains: A Case Study in Hindi Tourism QA
- 用大模型生成合成问答对,扩充有限的真实数据
- 小模型经多阶段微调后在印地语旅游任务上表现显著提升
- 适合低资源语言领域的小模型适配,可推广至其他垂直场景
低资源语言的领域特定问答面临两大挑战:标注数据稀缺和通用语言模型缺乏领域知识。本文提出一种多阶段微调策略,利用原始数据与合成数据,将轻量级语言模型适配至印地语旅游领域。合成问答对由大模型(LLaMA-70B、Phi-14B)生成,用于增强有限的原始数据集。我们探索多种训练方法并分析其对领域泛化能力的影响。结果表明,大模型能高效生成高质量合成数据,小模型可有效学习并适应这些数据,为低资源、领域特定的问答提供可扩展的解决方案。
原文摘要 · Abstract (English)
Domain-specific question answering in low-resource languages faces two key challenges: scarcity of annotated datasets and limited domain knowledge in general-purpose language models. In this work, we present a multi-stage finetuning strategy to adapt lightweight language models to the Hindi tourism domain by leveraging both original and synthetic training data. Synthetic question-answer pairs are generated using large LLMs (LLaMA-70B, Phi-14B) and used to augment the limited original dataset. We explore several training methodologies and analyse their impact on domain generalisation. Our results demonstrate that large models can efficiently generate synthetic data, while small models can effectively adapt to it, offering a scalable pathway for low-resource, domain-specific QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。