arXiv:2509.00092cs.LGcs.AI2025-09AAAI被引 1

提出新模型检测真实世界中结构多变的合成表格数据。

Robust Detection of Synthetic Tabular Data under Schema Variability

  • 设计逐条处理的Transformer架构,适应不同表格结构。
  • 在未知模式下准确率提升7个百分点,AUC与准确率均超基线。
  • 适合需要高鲁棒性检测合成数据的工业应用者。

生成模型的兴起引发了对数据真实性的担忧。尽管图像和文本的合成数据检测方法已广泛研究,但作为广泛应用的数据形式,表格数据的检测仍被忽视。由于表格结构异质且测试时可能出现未见过的格式,检测合成表格数据尤为困难。本文解决在真实场景下(即检测器部署于结构可变且此前未见的表格)检测合成表格数据的挑战。提出一种新型逐条处理的Transformer架构,显著优于唯一先前发布的基线,在AUC和准确率上分别提升7个百分点。通过引入表格自适应组件,模型准确率再增7点,展现出更强鲁棒性。本工作首次提供强有力证据,证明在现实条件下检测合成表格数据是可行的,并大幅超越现有方法。论文接受后,正完成代码开源的行政与授权流程,更新版本将在发布后立即补充。

原文摘要 · Abstract (English)

The rise of powerful generative models has sparked concerns over data authenticity. While detection methods have been extensively developed for images and text, the case of tabular data, despite its ubiquity, has been largely overlooked. Yet, detecting synthetic tabular data is especially challenging due to its heterogeneous structure and unseen formats at test time. We address the underexplored task of detecting synthetic tabular data ''in the wild'', i.e. when the detector is deployed on tables with variable and previously unseen schemas. We introduce a novel datum-wise transformer architecture that significantly outperforms the only previously published baseline, improving both AUC and accuracy by 7 points. By incorporating a table-adaptation component, our model gains an additional 7 accuracy points, demonstrating enhanced robustness. This work provides the first strong evidence that detecting synthetic tabular data in real-world conditions is feasible, and demonstrates substantial improvements over previous approaches. Following acceptance of the paper, we are finalizing the administrative and licensing procedures necessary for releasing the source code. This extended version will be updated as soon as the release is complete.

表格生成数据检测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。