用合成数据提升多语言农业问答模型的准确性和本地化水平
Leveraging Synthetic Data for Question Answering with Multilingual LLMs in the Agricultural Domain
- 从印度农业文档生成英、印地、旁遮普语合成数据
- 微调后模型在事实性、相关性上显著优于基线
- 适合需要多语言农业信息支持的开发者与研究者
让农民能及时获取母语的精准农业信息对农业成功至关重要。现有的通用大语言模型通常提供泛化建议,缺乏本地化和多语言场景下的准确性。本研究通过从印度农业特定文档生成英文、印地语、旁遮普语的合成数据集,并对大语言模型进行问答任务微调。在人工构建的数据集上评估显示,微调后的模型在事实性、相关性和农业共识方面均显著优于基线模型。
原文摘要 · Abstract (English)
Enabling farmers to access accurate agriculture-related information in their native languages in a timely manner is crucial for the success of the agriculture field. Publicly available general-purpose Large Language Models (LLMs) typically offer generic agriculture advisories, lacking precision in local and multilingual contexts. Our study addresses this limitation by generating multilingual (English, Hindi, Punjabi) synthetic datasets from agriculture-specific documents from India and fine-tuning LLMs for the task of question answering (QA). Evaluation on human-created datasets demonstrates significant improvements in factuality, relevance, and agricultural consensus for the fine-tuned LLMs compared to the baseline counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。