arXiv:2603.19712cs.CL2026-03

用多视角差异检测AI伪造的科研表格,准确率超98%。

TAB-AUDIT: Detecting AI-Fabricated Scientific Tables via Multi-View Likelihood Mismatch

  • 通过表格骨架与数值间的困惑度差距识别伪造痕迹
  • 在域内检测达0.987 AUROC,域外也达0.883 AUROC
  • 首个专门针对伪造表格的基准数据集,适合论文审查者使用

AI生成的虚假科学论文引发学术诚信危机。本文首次系统研究在实证NLP论文中检测AI伪造表格的方法,因表格内容是支撑论点的关键证据。我们构建了首个包含1,173篇AI生成和1,215篇人工撰写论文的基准数据集FabTab。通过分析发现,伪造表格与真实表格存在系统性差异,并提炼出一套可判别特征,形成TAB-AUDIT框架。核心特征“表内不匹配”捕捉表格骨架与数值内容之间的困惑度差异。实验表明,基于这些特征的随机森林模型显著优于现有方法,在域内达到0.987 AUROC,域外为0.883 AUROC。研究揭示实验表格是检测AI伪造的重要法医信号,并为后续研究提供新基准。

原文摘要 · Abstract (English)

AI-generated fabricated scientific manuscripts raise growing concerns with large-scale breaches of academic integrity. In this work, we present the first systematic study on detecting AI-generated fabricated scientific tables in empirical NLP papers, as information in tables serve as critical evidence for claims. We construct FabTab, the first benchmark dataset of fabricated manuscripts with tables, comprising 1,173 AI-generated papers and 1,215 human-authored ones in empirical NLP. Through a comprehensive analysis, we identify systematic differences between fabricated and real tables and operationalize them into a set of discriminative features within the TAB-AUDIT framework. The key feature, within-table mismatch, captures the perplexity gap between a table's skeleton and its numerical content. Experimental results show that RandomForest built on these features significantly outperform prior state-of-the-art methods, achieving 0.987 AUROC in-domain and 0.883 AUROC out-of-domain. Our findings highlight experimental tables as a critical forensic signal for detecting AI-generated scientific fraud and provide a new benchmark for future research.

AI检测科研诚信表格生成法医分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。