用635个真实与合成数据验证论文与代码一致性,发现大模型仅能识别46.7%真实差异。
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
- 构建跨模态验证数据集SciCoQA,涵盖真实与合成的论文-代码不一致问题。
- 最先进大模型仅检测到46.7%的真实差异,尤其在长上下文和缺失细节时表现差。
- 适用于关注科研可复现性、自动化质量评估的研究者与开发者。
科学论文与代码之间的不一致严重威胁研究可复现性,而随着自动化研究代理的扩张,人工审核已难以为继。目前尚无系统性评估大语言模型(LLM)识别此类不一致的能力。为此,我们提出SciCoQA,一个包含635个论文-代码不一致样本的数据集(92个真实,543个合成),用于跨模态验证任务。在22个评估模型中,即使表现最佳的Gemini 3.1 Pro和GPT-5 Mini也仅能检测到46.7%的真实差异,暴露出自动化科学质量保证的重大短板。SciCoQA源自GitHub问题与可复现性论文,并设计了合成生成流程,以扩展至物理、定量生物学等计算科学领域。我们进一步提出了差异类型分类体系,用于刻画不一致模式。分析表明,模型在处理遗漏论文细节、长上下文输入以及训练语料外的论文时尤为困难。
原文摘要 · Abstract (English)
Discrepancies between scientific papers and their code undermine reproducibility, a concern that grows as automated research agents scale scientific output beyond human review capacity. Whether LLMs can reliably detect such discrepancies has not been systematically measured. To this end, we present SciCoQA, a dataset of 635 paper-code discrepancies (92 real, 543 synthetic) for this cross-modal verification task. Across 22 evaluated models, even the best-performing LLMs, Gemini 3.1 Pro and GPT-5 Mini, detect only 46.7% of real-world discrepancies, revealing a critical gap in automated scientific quality assurance. We construct SciCoQA from GitHub issues and reproducibility papers, and propose a synthetic generation pipeline to scale beyond AI to Physics, Quantitative Biology, and other computational sciences. We further introduce a taxonomy of discrepancy types and categories to characterize the occurring mismatches. Our analysis shows that models particularly struggle with omitted paper details, long-context inputs, and papers outside their pre-training corpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。