arXiv:2603.29791cs.AIcs.CL2026-03中稿 · TMLR 2026, J2C Cer…被引 4

无需种子数据,用推理驱动生成可解释的合成数据

Reasoning-Driven Synthetic Data Generation and Evaluation

  • 通过无种子代理机制,让用户以可控方式定义数据特征
  • 在多个数据集上验证了生成数据的内在与下游性能
  • 适合数据稀缺或隐私敏感领域的AI开发应用

许多感兴趣的AI应用需要专门的多模态模型,但训练此类模型的相关数据往往稀缺或不可获取。依赖人工标注填补这些缺口成本过高、易出错且耗时,促使模型开发者越来越多地考虑合成数据作为可扩展的替代方案。然而,现有合成数据生成方法通常依赖手动提示、进化算法或目标分布的大量种子数据,限制了其可扩展性、可解释性和控制力。本文提出Simula:一种新颖的推理驱动型数据生成与评估框架。该框架采用无种子、代理式方法,在大规模生成合成数据集的同时,允许用户通过可解释且可控的过程定义期望的数据特征,实现细粒度资源分配。我们在多种数据集上验证了该方法的有效性,严格测试了数据的内在属性和下游表现。本工作(1)提供合成数据机制设计指南,(2)揭示大规模生成与评估合成数据的洞见,(3)为数据稀缺或隐私敏感领域中的AI开发与部署开辟新可能。

原文摘要 · Abstract (English)

Although many AI applications of interest require specialized multi-modal models, relevant data to train such models is inherently scarce or inaccessible. Filling these gaps with human annotators is prohibitively expensive, error-prone, and time-consuming, leading model builders to increasingly consider synthetic data as a scalable alternative. However, existing synthetic data generation methods often rely on manual prompts, evolutionary algorithms, or extensive seed data from the target distribution - limiting their scalability, explainability, and control. In this paper, we introduce Simula: a novel reasoning-driven framework for data generation and evaluation. It employs a seedless, agentic approach to generate synthetic datasets at scale, allowing users to define desired dataset characteristics through an explainable and controllable process that enables fine-grained resource allocation. We show the efficacy of our approach on a variety of datasets, rigorously testing both intrinsic and downstream properties. Our work (1) offers guidelines for synthetic data mechanism design, (2) provides insights into generating and evaluating synthetic data at scale, and (3) unlocks new opportunities for developing and deploying AI in domains where data scarcity or privacy concerns are paramount.

合成数据推理驱动多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。