arXiv:2602.04785cs.LGcs.AI2026-02被引 1

用多模型协作生成高质量表格数据,再三重质检。

Team, Then Trim: An Assembly-Line LLM Framework for High-Quality Tabular Data Generation

  • 多个大模型分工生成表格不同部分,像流水线生产。
  • 在真实和模拟数据上均优于现有方法,提升数据质量。
  • 适合数据稀缺场景,如医疗、金融等低样本领域。

表格数据是众多机器学习应用的基础,但获取高质量表格数据通常耗时且昂贵。受限于观测数量不足,表格数据常存在类别不平衡、选择偏差和低保真度等严重缺陷。针对这些问题,本文基于大语言模型(LLM)的最新进展,提出团队-修剪(Team-then-Trim, T²)框架:通过一组协作的LLM生成合成表格数据,并引入严格的三阶段插件式数据质量控制(QC)流程。在T²中,表格生成被构想为制造过程:由领域知识引导的专用LLM按序生成数据的不同组件,生成结果——即合成数据——在多个质量维度上被系统评估。在模拟和真实数据集上的实证结果表明,T²在生成高质量表格数据方面显著优于当前最优方法,展现出在直接数据收集不可行时支持下游模型的巨大潜力。

原文摘要 · Abstract (English)

While tabular data is fundamental to many real-world machine learning (ML) applications, acquiring high-quality tabular data is usually labor-intensive and expensive. Limited by the scarcity of observations, tabular datasets often exhibit critical deficiencies, such as class imbalance, selection bias, and low fidelity. To address these challenges, building on recent advances in Large Language Models (LLMs), this paper introduces Team-then-Trim (T$^2$), a framework that synthesizes high-quality tabular data through a collaborative team of LLMs, followed by a rigorous three-stage plug-in data quality control (QC) pipeline. In T$^2$, tabular data generation is conceptualized as a manufacturing process: specialized LLMs, guided by domain knowledge, are tasked with generating different data components sequentially, and the resulting products, i.e., the synthetic data, are systematically evaluated across multiple dimensions of QC. Empirical results on both simulated and real-world datasets demonstrate that T$^2$ outperforms state-of-the-art methods in producing high-quality tabular data, highlighting its potential to support downstream models when direct data collection is practically infeasible.

表格生成大模型数据合成质量控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。