提出四类无需表格先验的合成表格检测方法
Synthetic Tabular Data Detection In the Wild
- 设计不依赖表格结构的通用检测器,配合简单预处理
- 在六种不同复杂度场景下验证,跨表迁移仍具挑战
- 适合数据安全、可信分析领域研究人员参考
检测合成表格数据对防止虚假或篡改数据传播至关重要,可能影响基于数据的决策。本研究探讨了在不同表格间是否能可靠识别合成数据。这一挑战独特于表格数据,因表格结构(如列数、数据类型、格式)差异大。我们提出四种不依赖表格的检测器,结合简单预处理方案,在六个评估协议下进行测试,涵盖不同程度的‘野外’复杂性。结果表明,在有限表格集合上实现跨表学习是可行的,即使采用基础预处理。然而,跨表迁移(即部署到未见过的表格)仍具挑战,暗示需更复杂的编码方案来应对该问题。
原文摘要 · Abstract (English)
Detecting synthetic tabular data is essential to prevent the distribution of false or manipulated datasets that could compromise data-driven decision-making. This study explores whether synthetic tabular data can be reliably identified across different tables. This challenge is unique to tabular data, where structures (such as number of columns, data types, and formats) can vary widely from one table to another. We propose four table-agnostic detectors combined with simple preprocessing schemes that we evaluate on six evaluation protocols, with different levels of ''wildness''. Our results show that cross-table learning on a restricted set of tables is possible even with naive preprocessing schemes. They confirm however that cross-table transfer (i.e. deployment on a table that has not been seen before) is challenging. This suggests that sophisticated encoding schemes are required to handle this problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。