arXiv:2505.23852cs.CLcs.AI2025-05

用大模型模拟研究团队,自动复现阿尔茨海默病研究结果。

Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease

  • 构建基于GPT-4o的智能代理团队,仅凭摘要和方法复现研究。
  • 平均复现率53.2%,数值与原结果常有差异。
  • 适合关注科研可复现性评估的学者参考。

目标:展示大型语言模型(LLMs)作为自主代理,利用相同或相似数据集复现已发表研究的能力。材料与方法:使用国家阿尔茨海默病协调中心(NACC)的“快速访问”数据集,筛选出五篇高被引、可复现的研究。基于GPT-4o,创建由多个自主代理组成的模拟研究团队,任务是根据论文摘要、方法部分及数据字典描述,动态编写并执行代码复现研究结果。结果:共提取5项研究中的35个关键发现。平均而言,每个研究的发现被约53.2%成功复现。数值结果和范围类结论在不同研究与代理间存在差异,统计方法或参数也常不一致,但整体趋势与显著性有时仍相似。讨论:部分案例中代理成功复制了研究技术与结果,另一些则因实现缺陷或方法细节缺失而失败。这些差异揭示了当前大模型在完全自动化复现评估方面的局限性。然而,本初步研究凸显了结构化代理系统在规模化评估科研严谨性方面的潜力。结论:该探索性工作展示了大模型作为自主代理在生物医学研究可复现性自动化中的前景与不足。

原文摘要 · Abstract (English)

Objective: To demonstrate the capabilities of Large Language Models (LLMs) as autonomous agents to reproduce findings of published research studies using the same or similar dataset. Materials and Methods: We used the "Quick Access" dataset of the National Alzheimer's Coordinating Center (NACC). We identified highly cited published research manuscripts using NACC data and selected five studies that appeared reproducible using this dataset alone. Using GPT-4o, we created a simulated research team of LLM-based autonomous agents tasked with writing and executing code to dynamically reproduce the findings of each study, given only study Abstracts, Methods sections, and data dictionary descriptions of the dataset. Results: We extracted 35 key findings described in the Abstracts across 5 Alzheimer's studies. On average, LLM agents approximately reproduced 53.2% of findings per study. Numeric values and range-based findings often differed between studies and agents. The agents also applied statistical methods or parameters that varied from the originals, though overall trends and significance were sometimes similar. Discussion: In some cases, LLM-based agents replicated research techniques and findings. In others, they failed due to implementation flaws or missing methodological detail. These discrepancies show the current limits of LLMs in fully automating reproducibility assessments. Still, this early investigation highlights the potential of structured agent-based systems to provide scalable evaluation of scientific rigor. Conclusion: This exploratory work illustrates both the promise and limitations of LLMs as autonomous agents for automating reproducibility in biomedical research.

大模型科研复现阿尔茨海默病智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。