arXiv:2605.08138cs.LG2026-05ACL被引 1

一站式合成数据工具,支持多模态多语言高效生成

DataArc-SynData-Toolkit: A Unified Closed-Loop Framework for Multi-Path, Multimodal, and Multilingual Data Synthesis

论文配图:DataArc-SynData-Toolkit: A Unified Closed-Loop Framework for Multi-Path, Multimodal, and Multilingual Data Synthesis
图 1 · 摘自论文原文
  • 配置驱动的全流程闭环框架,可视化操作降低使用门槛
  • 统一标准实现高质量多源数据合成,提升可复用性
  • 模块化设计适配多任务、多语言场景,适合研究与工程落地

合成数据已成为解决大语言模型在特定领域和低资源语言中数据稀缺问题的关键方案。然而,现有合成数据工具因流程复杂、标准分散、跨模态扩展性差而难以推广。为此,我们开发了开源的DataArc-SynData-Toolkit,具备:(1) 配置驱动的端到端流水线,配备直观的可视化界面和简化命令行,显著提升易用性;(2) 统一且质量可控的合成范式,标准化多源数据生成,确保高可复用性;(3) 高度模块化架构,支持无缝适配多模态、多语言和多任务需求。我们在多个应用场景中验证该工具,实验表明其在生成效率与数据质量间达到最优平衡。通过提供端到端、可视化交互的流程,DataArc-SynData-Toolkit大幅降低合成数据生成及后续模型训练的技术门槛,加速其在实际应用中的部署。

原文摘要 · Abstract (English)

Synthetic data has emerged as a crucial solution to the data scarcity bottleneck in large language models (LLMs), particularly for specialized domains and low-resource languages. However, the broader adoption of existing synthetic data tools is severely hindered by convoluted workflows, fragmented data standards, and limited scalability across modalities. To address these limitations, we develop DataArc-SynData-Toolkit, an open-source framework featuring: (1) a configuration-driven, end-to-end pipeline equipped with an intuitive visual interface and simplified CLI for exceptional usability; (2) a unified, quality-controllable synthesis paradigm that standardizes multi-source data generation to ensure high reusability; and (3) a highly modular architecture designed for seamless multimodal, multilingual, and multi-task adaptation. We apply the toolkit in multiple application scenarios. Experimental results demonstrate that our toolkit achieves an optimal balance between generation efficiency and data quality. By offering an end-to-end and visually interactive pipeline, DataArc-SynData-Toolkit significantly lowers the technical barrier to synthetic data generation and subsequent model training, accelerating its practical deployment in real-world applications.

合成数据多模态多语言LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。