arXiv:2504.20900cs.LG2025-04被引 2

提出三项新指标,更全面评估表格数据生成模型性能。

Evaluating Generative Models for Tabular Data: Novel Metrics and Benchmarking

  • 设计三类新指标:FAED、FPCAD、RFIS,捕捉传统方法忽略的生成问题。
  • 在三个入侵检测数据集上验证,FAED能有效识别其他指标遗漏的问题。
  • 适合研究表格生成模型评估的学者和工业界从业者参考使用。

生成模型已在多个领域带来变革,但其在表格数据上的应用仍不充分。由于结构复杂、规模差异大及混合数据类型,表格生成模型的评估面临独特挑战,现有指标难以直观捕捉其复杂模式。当前评价方法仅提供部分信息,缺乏对生成性能的全面衡量。为此,本文提出三种新型评估指标:FAED、FPCAD 和 RFIS。我们在三个标准网络入侵检测数据集上进行了广泛实验,将这些指标与现有的 Fidelity、Utility、TSTR、TRTS 方法进行对比。结果表明,FAED 能有效识别出其他指标未能捕捉的生成建模问题;而 FPCAD 表现有潜力,但仍需进一步优化以提高可靠性。所提出的框架为表格数据生成模型的评估提供了稳健且实用的方法。

原文摘要 · Abstract (English)

Generative models have revolutionized multiple domains, yet their application to tabular data remains underexplored. Evaluating generative models for tabular data presents unique challenges due to structural complexity, large-scale variability, and mixed data types, making it difficult to intuitively capture intricate patterns. Existing evaluation metrics offer only partial insights, lacking a comprehensive measure of generative performance. To address this limitation, we propose three novel evaluation metrics: FAED, FPCAD, and RFIS. Our extensive experimental analysis, conducted on three standard network intrusion detection datasets, compares these metrics with established evaluation methods such as Fidelity, Utility, TSTR, and TRTS. Our results demonstrate that FAED effectively captures generative modeling issues overlooked by existing metrics. While FPCAD exhibits promising performance, further refinements are necessary to enhance its reliability. Our proposed framework provides a robust and practical approach for assessing generative models in tabular data applications.

生成模型表格数据评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。