Data-Juicer 2.0实现海量多模态数据的自适应处理,支持千亿级数据高效训练。
Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
- 集成100+多模态数据处理算子,支持文本、图像、视频、音频全链路处理
- 可调度超1万核CPU,高效处理TB级数据,实测性能显著优于前代
- 提供可视化界面与对话式交互,适合科研与工业界快速部署基础模型
基础模型需要对大规模多模态数据进行高级处理,但传统框架难以应对多模态数据的独特复杂性。为此,我们提出 Data-Juicer 2.0,一个包含100多个跨文本、图像、视频、音频模态的数据处理算子的系统,支持数据分析、合成、标注及基础模型后训练等关键任务。该系统与 Hugging Face、Ray 等主流数据集平台和计算引擎无缝兼容,并通过用户友好的接口层(支持 Python、RESTful API 和对话命令)提升易用性与可编程性。其新运行时层具备跨规模与环境的自适应执行能力,抽象底层系统复杂性。大量实证评估表明,Data-Juicer 2.0 在性能与可扩展性方面表现卓越,可在超过1万核的CPU上高效处理TB级数据。系统已开源,被广泛应用于阿里巴巴云 PAI 等研究与产品场景。我们持续维护并分享实践经验,推动下一代基础模型的研究与应用。
原文摘要 · Abstract (English)
Foundation models demand advanced data processing for their vast, multimodal datasets. However, traditional frameworks struggle with the unique complexities of multimodal data. In response, we present Data-Juicer 2.0, a data processing system backed by 100+ data processing operators spanning text, image, video, and audio modalities, supporting more critical tasks including data analysis, synthesis, annotation, and foundation model post-training. With seamless compatibility and dedicated optimization for popular dataset hubs like Hugging Face and computing engines like Ray, it improves upon its predecessor in terms of usability, efficiency, and programmability. It features an easily accessible user interface layer that supports decoupled Python interactions, RESTful APIs, and conversational commands. Its new runtime layer offers adaptive execution across diverse scales and environments, abstracting away system complexities. Extensive empirical evaluations demonstrate Data-Juicer 2.0's remarkable performance and scalability, highlighting its capability to efficiently process TB-level data with 10k+ CPU cores. The system is publicly available and has been widely adopted in diverse research fields and real-world products such as Alibaba Cloud PAI. We actively maintain the system and share practical insights to foster research and applications of next-generation foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。