一站式生成高质量对话数据,自动打标管理,省去人工调优
SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
- 模块化流程+配置驱动,轻松构建复杂对话流
- 双阶段质量评分:规则+大模型,自动筛选高质对话
- 支持SFT与DPO,适合大规模语言模型训练场景
大语言模型的进展高度依赖于高质量数据集,用于监督微调(SFT)和直接偏好优化(DPO)等对齐任务。本文提出一个统一的合成数据生成框架,支持可扩展、可配置、高保真的合成数据生成,专为上述训练范式设计。该框架采用模块化、配置驱动的流水线,能以最少的人工干预建模复杂对话流程。通过双阶段质量打标机制——结合启发式规则与大模型评估——自动过滤并评分来自OASST格式对话的数据,确保高质量对话样本的选取。生成的数据遵循灵活的结构化模式,同时支持SFT与DPO应用,便于无缝集成到多种训练流程中。整体方案实现大规模合成对话数据的高效生成与管理,显著降低大模型训练中的数据准备成本。
原文摘要 · Abstract (English)
The advancement of large language models (LLMs) is critically dependent on the availability of high-quality datasets for Supervised Fine-Tuning (SFT), alignment tasks like Direct Preference Optimization (DPO), etc. In this work, we present a comprehensive synthetic data generation framework that facilitates scalable, configurable, and high-fidelity generation of synthetic data tailored for these training paradigms. Our approach employs a modular and configuration-based pipeline capable of modeling complex dialogue flows with minimal manual intervention. This framework uses a dual-stage quality tagging mechanism, combining heuristic rules and LLM-based evaluations, to automatically filter and score data extracted from OASST-formatted conversations, ensuring the curation of high-quality dialogue samples. The resulting datasets are structured under a flexible schema supporting both SFT and DPO use cases, enabling seamless integration into diverse training workflows. Together, these innovations offer a robust solution for generating and managing synthetic conversational data at scale, significantly reducing the overhead of data preparation in LLM training pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。