评估合成表格数据中列间逻辑关系的保真度,填补了生成质量评估的空白。
Evaluating Inter-Column Logical Relationships in Synthetic Tabular Data Generation
- 提出三种衡量列间逻辑关系保真度的评估指标
- 发现现有方法难以保持地理层级、时间序列等关键逻辑关系
- 适合关注数据真实性和可用性的研究人员与工业应用者
当前合成表格数据的评估主要关注联合分布建模效果,常忽视对事件序列和跨列实体关系一致性的评估。本文提出三项评估指标,用于衡量合成表格数据中列间逻辑关系的保留程度。通过在真实工业数据集上测试经典与前沿生成方法,实验表明现有方法在维持地理或组织层级等层次关系、时间序列及数学依赖关系方面普遍表现不佳,这些关系对真实表格数据的细粒度真实性至关重要。基于此,研究进一步探讨了在建模分布的同时更好捕捉逻辑关系的可行路径。代码已公开于 https://github.com/Yunbo-max/TabLogicEval。
原文摘要 · Abstract (English)
Current evaluations of synthetic tabular data mainly focus on how well joint distributions are modeled, often overlooking the assessment of their effectiveness in preserving realistic event sequences and coherent entity relationships across columns.This paper proposes three evaluation metrics designed to assess the preservation of logical relationships among columns in synthetic tabular data. We validate these metrics by assessing the performance of both classical and state-of-the-art generation methods on a real-world industrial dataset.Experimental results reveal that existing methods often fail to rigorously maintain logical consistency (e.g., hierarchical relationships in geography or organization) and dependencies (e.g., temporal sequences or mathematical relationships), which are crucial for preserving the fine-grained realism of real-world tabular data. Building on these insights, this study also discusses possible pathways to better capture logical relationships while modeling the distribution of synthetic tabular data. The code is available at https://github.com/Yunbo-max/TabLogicEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。