arXiv:2412.13227cs.LGcs.DB2024-12

提出跨表格合成数据检测方法,应对真实场景中多样生成器与格式的挑战。

Cross-table Synthetic Tabular Data Detection

  • 设计三种跨表格基准检测器,适应不同生成器与格式
  • 在四种不同真实度评估协议下验证,检测准确率低于60%
  • 为防范伪造数据传播提供新方向,适合数据安全研究者

检测合成表格数据对于防止虚假或被篡改的数据集传播至关重要,以免影响数据驱动的决策。本研究探讨合成表格数据是否能在真实环境中可靠识别——即跨越不同生成器、领域和表格格式。这一挑战对表格数据尤为独特,因其结构(如列数、数据类型、格式)差异极大。我们提出三种跨表格基准检测器和四种不同的评估协议,分别对应不同程度的‘真实感’。初步结果表明,跨表格适应仍是一项极具挑战的任务。

原文摘要 · Abstract (English)

Detecting synthetic tabular data is essential to prevent the distribution of false or manipulated datasets that could compromise data-driven decision-making. This study explores whether synthetic tabular data can be reliably identified ''in the wild''-meaning across different generators, domains, and table formats. This challenge is unique to tabular data, where structures (such as number of columns, data types, and formats) can vary widely from one table to another. We propose three cross-table baseline detectors and four distinct evaluation protocols, each corresponding to a different level of ''wildness''. Our very preliminary results confirm that cross-table adaptation is a challenging task.

数据检测合成数据表格数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。