arXiv:2410.08800cs.CL2024-10被引 3

构建多语言大模型的数据流水线,聚焦欧盟实际应用。

Data Processing for the OpenGPT-X Model Family

  • 区分精选数据与网络数据,分别设计处理流程。
  • 通过去重和过滤生成高质量训练数据集。
  • 适合关注多语言模型与合规数据处理的研究者。

本文全面介绍了OpenGPT-X项目所开发的数据准备流程,该项目旨在创建开放且高性能的多语言大型语言模型(LLMs),覆盖所有主要欧洲语言,并特别关注在欧盟范围内的实际应用。从数据选择与需求定义到最终筛选数据的准备,详细说明了整个数据处理步骤。区分了精选数据与网络数据,前者经过少量过滤,后者需进行大量过滤与去重处理,由此催生了针对两类数据的专用算法方案。除了描述处理方法外,还对数据集进行了深入分析,提升透明度并符合欧洲数据法规要求。最后分享了项目中的关键洞察与挑战,为未来大规模多语言数据准备提供参考。

原文摘要 · Abstract (English)

This paper presents a comprehensive overview of the data preparation pipeline developed for the OpenGPT-X project, a large-scale initiative aimed at creating open and high-performance multilingual large language models (LLMs). The project goal is to deliver models that cover all major European languages, with a particular focus on real-world applications within the European Union. We explain all data processing steps, starting with the data selection and requirement definition to the preparation of the final filtered data. We distinguish between curated data and web data, as each of these categories is handled by distinct pipelines, with curated data undergoing minimal filtering and web data requiring extensive filtering and deduplication. This distinction guided the development of specialized algorithmic solutions for both pipelines. In addition to describing the processing methodologies, we provide an in-depth analysis of the datasets, increasing transparency and alignment with European data regulations. Finally, we share key insights and challenges faced during the project, offering recommendations for future endeavors in large-scale multilingual data preparation for LLMs.

多语言模型数据处理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。