arXiv:2504.08829cs.LGcs.AI2025-04

提出新模型检测真实世界中未知结构的合成表格数据

Datum-wise Transformer for Synthetic Tabular Data Detection in the Wild

  • 按记录维度设计新型Transformer,适应多样表格结构
  • 在未见表格上检测准确率显著优于现有方法
  • 结合领域自适应提升泛化能力,适合实际部署场景

生成模型能力的提升引发对发表内容真实性的担忧。已有检测方法多针对图像或文本等结构化媒体,但对工业与政府领域重要的合成表格数据检测研究较少。由于表格结构差异大(列数、类型变化剧烈),该任务极具挑战性。本文解决真实世界中“未知结构”表格的合成数据检测问题,提出一种全新的基于记录维度的Transformer架构,实验表明其性能优于现有模型。同时引入领域自适应技术增强模型鲁棒性,为数据伪造检测提供更可靠的解决方案。

原文摘要 · Abstract (English)

The growing power of generative models raises major concerns about the authenticity of published content. To address this problem, several synthetic content detection methods have been proposed for uniformly structured media such as image or text. However, little work has been done on the detection of synthetic tabular data, despite its importance in industry and government. This form of data is complex to handle due to the diversity of its structures: the number and types of the columns may vary wildly from one table to another. We tackle the tough problem of detecting synthetic tabular data ''in the wild'', i.e. when the model is deployed on table structures it has never seen before. We introduce a novel datum-wise transformer architecture and show that it outperforms existing models. Furthermore, we investigate the application of domain adaptation techniques to enhance the effectiveness of our model, thereby providing a more robust data-forgery detection solution.

表格生成伪造检测Transformer领域自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。