用大模型检测医学论文造假,还能给出可信解释。
Pub-Guard-LLM: Detecting Retracted Biomedical Articles with Reliable Explanations
- 基于大模型设计三种检测模式,支持可解释预测。
- 在超1.1万篇真实医学文章上表现优于基线。
- 开源数据集助力科研诚信,适合研究者与期刊使用。
大量已发表的科学论文被发现存在欺诈行为,严重威胁医学等领域的研究可信度与安全性。我们提出Pub-Guard-LLM,首个专用于生物医学论文欺诈检测的大语言模型系统。提供三种部署模式:原始推理、检索增强生成、多智能体辩论,均支持文本化预测解释。为评估性能,我们构建了开源基准PubMed Retraction,包含超过11,000篇真实生物医学文章及重申标签。实验表明,各模式下Pub-Guard-LLM均显著优于多种基线方法,且生成的解释在相关性与连贯性上更优,经多重评估验证。该系统提升了科学欺诈检测的准确性与可解释性,为维护科研诚信提供了新颖、高效、开源的工具。
原文摘要 · Abstract (English)
A significant and growing number of published scientific articles is found to involve fraudulent practices, posing a serious threat to the credibility and safety of research in fields such as medicine. We propose Pub-Guard-LLM, the first large language model-based system tailored to fraud detection of biomedical scientific articles. We provide three application modes for deploying Pub-Guard-LLM: vanilla reasoning, retrieval-augmented generation, and multi-agent debate. Each mode allows for textual explanations of predictions. To assess the performance of our system, we introduce an open-source benchmark, PubMed Retraction, comprising over 11K real-world biomedical articles, including metadata and retraction labels. We show that, across all modes, Pub-Guard-LLM consistently surpasses the performance of various baselines and provides more reliable explanations, namely explanations which are deemed more relevant and coherent than those generated by the baselines when evaluated by multiple assessment methods. By enhancing both detection performance and explainability in scientific fraud detection, Pub-Guard-LLM contributes to safeguarding research integrity with a novel, effective, open-source tool.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。