arXiv:2608.15109cs.AI2026-08

用大模型发现并修复表格数据的列间约束,让合成数据既真实又合规。

Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents

论文配图:Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents
图 1 · 摘自论文原文
  • 用大模型自动发现方程、不等式和逻辑依赖三类列间约束
  • 修复后合成数据零违规,且保留单变量分布特征
  • 适用于需要数据合规性的金融、医疗等领域

生成结构有效的合成表格数据仍具挑战:高统计保真度的输出可能违反有意义的语义约束。本文研究三类互补的列间约束——方程、线性不等式与逻辑依赖的发现与强制执行。提出的统一工具化流程将三类约束均表示为可机器执行的假设,并通过统一接口实现全表验证、确定性诊断与反例引导修正。一个与生成器无关的后处理器对各类约束进行针对性修复,保持原始生成器不变。在精心设计的行为审计与端到端评估中,完整流程在未见约束检测上优于一次性直接提示方法;后处理使所有适用约束下的测量违规数为零,多数数据集下游任务性能提升,且基本保持单变量边缘分布。

原文摘要 · Abstract (English)

Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical fidelity and downstream utility can still violate semantically meaningful domain constraints. We study the discovery and enforcement of three complementary inter-column constraint families---equations, linear inequalities, and logical dependencies. Our unified tool-grounded workflow represents all three as machine-executable hypotheses and applies a common interface for full-table validation, deterministic diagnosis, and counterexample-guided revision. A generator-agnostic postprocessor coordinates family-specific repairs on outputs from unchanged tabular generators. Across curated behavioral audits and end-to-end evaluations, the complete workflow improves held-out violation detection over one-shot direct prompting, while postprocessing yields zero measured violations for every retained, applicable constraint, improves downstream utility on most datasets, and largely preserves univariate marginals.

表格生成约束修复大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。