用机器学习检测科研数据中人为造假的数字规律异常。
A machine-learning-assisted progressive digit-randomness screening framework for detecting non-random patterns in raw numerical research data
- 融合统计检验与机器学习,分析数据小数位模式是否随机。
- 在模拟造假数据中识别出显著异常,真实数据无明显问题。
- 可辅助筛查高风险数据,适合科研诚信审查人员使用。
原始数值数据在完整性审查中远不如图像、抄袭或统计摘要问题受重视。我们开发了伪造风险数字随机性筛查模型(FDRS),一种结合统计与机器学习的框架,用于检测科研数据中非随机的数字模式异常。FDRS整合了单/双小数位检验、Cramer's V、熵值、KL散度、数字偏好指数、渐进抽样和半监督风险评分。在仪器获取的酶吸收数据(RawData,n=253)和人工模拟异常数据(ErrData,n=255)上验证:RawData在第三位小数上无显著偏差,而ErrData存在显著异常;联合第三、四位小数分析显示,ErrData具有更高Cramer's V、更低归一化熵、更高KL散度及更持续的渐进抽样信号。内部验证中,弹性网络逻辑回归达到最高AUC(0.98395)和最低布里尔分数(0.048439),随机森林准确率最高(0.926667)且平衡准确率(0.935)。RawData获低集成风险分(0.124627),分类为0级;ErrData得高分(0.740760),分类为3级。外部实证支持分级风险分层:无公开质疑的三个数据集为0或1级,两例被公开质疑或机构介入的为2或3级。FDRS通过可解释特征整合,可优先筛选需进一步审查的原始数据,是辅助性数字结构筛查工具,不构成伪造或不当行为的独立证据。
原文摘要 · Abstract (English)
Raw numerical datasets remain less systematically examined in integrity screening than images, plagiarism, or summary-statistic inconsistencies. We developed the Fabrication-risk Digit Randomness Screening model (FDRS), a statistical and machine-learning framework for detecting non-random digit-pattern irregularities in numerical research data. FDRS integrates single- and joint-decimal-digit tests, Cramer's V, entropy metrics, Kullback-Leibler divergence, digit-preference indices, progressive subsampling, and semi-supervised risk scoring. It was evaluated using an instrument-derived enzymatic absorbance dataset (RawData, n=253) and a blinded manually simulated irregular dataset (ErrData, n=255). RawData showed no significant deviation in single third-decimal-digit analysis, whereas ErrData showed a significant deviation. In joint third-fourth decimal digit analysis, ErrData showed higher Cramer's V, lower normalized entropy, higher KL divergence, and a more persistent progressive-subsampling deviation signal. In internal validation, Elastic-net Logistic Regression achieved the highest AUC (0.98395) and lowest Brier score (0.048439), while Random Forest achieved the highest accuracy (0.926667) and balanced accuracy (0.935). RawData received a low ensemble risk score of 0.124627 and was classified as Grade 0; ErrData received a score of 0.740760 and was classified as Grade 3. External real-world benchmarks supported graded risk stratification: three datasets without identified public post-publication concerns were classified as Grade 0 or 1, whereas two datasets from publicly questioned or institutionally handled articles were classified as Grade 2 or 3. FDRS can prioritize raw numerical datasets for further review by integrating interpretable statistical and machine-learning features. It is an auxiliary digit-structure screening tool, not standalone evidence of fabrication or misconduct.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。