arXiv:2411.11289cs.CLcs.AI2024-11EMNLP被引 1

用CPU实现低成本高质数据流水线,让小团队也能训练专业大模型。

LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models

  • 全程在CPU运行,不用昂贵GPU即可完成数据清洗与筛选。
  • 相比传统流程,准备时间与成本大幅降低,数据质量仍保持高位。
  • 可定制特定领域和语言的数据集,适合垂直场景应用开发。

为大型语言模型(LLMs)构建高质量、大规模数据集通常依赖资源密集型的GPU加速模型进行质量过滤,导致过程耗时且成本高昂,限制了缺乏强大计算基础设施的组织参与。为此,我们提出轻量级、目标导向的(LP)数据流水线框架,完全基于CPU运行,实现了数据提取、过滤与编排的高效处理。基于四项核心原则,该流水线显著缩短了准备时间并降低了成本,同时保持了高数据质量。尤为重要的是,该方法支持创建针对特定领域与语言的目的导向数据集,提升了大模型在专业场景中的适用性。我们预期该流水线将降低大模型开发门槛,使更多组织能够便捷地获取与使用大模型。

原文摘要 · Abstract (English)

Creating high-quality, large-scale datasets for large language models (LLMs) often relies on resource-intensive, GPU-accelerated models for quality filtering, making the process time-consuming and costly. This dependence on GPUs limits accessibility for organizations lacking significant computational infrastructure. To address this issue, we introduce the Lightweight, Purpose-driven (LP) Data Pipeline, a framework that operates entirely on CPUs to streamline the processes of dataset extraction, filtering, and curation. Based on our four core principles, the LP Data Pipeline significantly reduces preparation time and cost while maintaining high data quality. Importantly, our pipeline enables the creation of purpose-driven datasets tailored to specific domains and languages, enhancing the applicability of LLMs in specialized contexts. We anticipate that our pipeline will lower the barriers to LLM development, enabling a wide range of organizations to access LLMs more easily.

数据流水线轻量化大模型低成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。