大模型能发现论文中的数据泄露问题,提升研究可信度。
Can Large Language Models Detect Methodological Flaws? Evidence from Gesture Recognition for UAV-Based Rescue Operation Based on Deep Learning

- 用大模型分析论文,自动识别评估协议中的数据泄露漏洞。
- 六种主流大模型均一致指出近满分准确率源于非独立数据划分。
- 适合科研人员和审稿人用于提升论文可复现性审查效率。
可靠评估对机器学习研究至关重要,但方法学缺陷——尤其是数据泄露——持续削弱已有结果的有效性。本文探究大型语言模型(LLMs)能否作为独立分析工具,识别已发表研究中的此类问题。以一篇在小型人工中心数据集上报告接近完美准确率的手势识别论文为例,我们首先证明其评估协议因训练与测试集非独立划分而存在主体级数据泄露。随后,评估六种先进LLM是否能在不依赖额外上下文的情况下,通过相同提示独立识别该缺陷。所有模型一致判定评估存在缺陷,并将报告性能归因于非独立数据分割,依据包括重叠的学习曲线、极小的泛化差距及接近完美的分类结果。这些发现表明,仅基于已发表成果,大模型即可检测常见方法学问题。虽非定论,但其高度一致性凸显其作为提升可复现性与支持科学审计的辅助工具潜力。
原文摘要 · Abstract (English)
Reliable evaluation is essential in machine learning research, yet methodological flaws-particularly data leakage-continue to undermine the validity of reported results. In this work, we investigate whether large language models (LLMs) can act as independent analytical agents capable of identifying such issues in published studies. As a case study, we analyze a gesture-recognition paper reporting near-perfect accuracy on a small, human-centered dataset. We first show that the evaluation protocol is consistent with subject-level data leakage due to non-independent training and test splits. We then assess whether this flaw can be detected independently by six state-of-the-art LLMs, each analyzing the original paper without prior context using an identical prompt. All models consistently identify the evaluation as flawed and attribute the reported performance to non-independent data partitioning, supported by indicators such as overlapping learning curves, minimal generalization gap, and near-perfect classification results. These findings suggest that LLMs can detect common methodological issues based solely on published artifacts. While not definitive, their consistent agreement highlights their potential as complementary tools for improving reproducibility and supporting scientific auditing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。