arXiv:2510.02241cs.IRcs.CL2025-10被引 1

30亿参数开源大模型可替代闭源模型生成训练数据

Study on LLMs for Promptagator-Style Dense Retriever Training

  • 用30亿参数以下的开源大模型生成特定任务查询
  • 小规模开源模型在密集检索微调中表现接近闭源模型
  • 适合资源有限或数据敏感场景的从业者参考

Promptagator表明,少量示例提示下的大型语言模型(LLMs)可用作领域专用密集检索模型微调的任务特定查询生成器。然而,原始的Promptagator方法依赖于专有且大规模的LLM,用户可能无法访问,或在处理敏感数据时受到限制。本文研究了可在合理规模(≤140亿参数)下使用的开源LLM作为替代方案的效果。结果表明,最小仅30亿参数的开源模型即可作为有效的Promptagator风格查询生成器。本工作旨在为缺乏访问权限或需保护数据安全的实践者提供可靠的合成数据生成方案,并为提升领域专用应用的微调效果提供洞见。

原文摘要 · Abstract (English)

Promptagator demonstrated that Large Language Models (LLMs) with few-shot prompts can be used as task-specific query generators for fine-tuning domain-specialized dense retrieval models. However, the original Promptagator approach relied on proprietary and large-scale LLMs which users may not have access to or may be prohibited from using with sensitive data. In this work, we study the impact of open-source LLMs at accessible scales ($\leq$14B parameters) as an alternative. Our results demonstrate that open-source LLMs as small as 3B parameters can serve as effective Promptagator-style query generators. We hope our work will inform practitioners with reliable alternatives for synthetic data generation and give insights to maximize fine-tuning results for domain-specific applications.

大模型检索增强开源模型数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。