arXiv:2601.11825cs.AIcs.IR2026-01

用AI自动合成医学研究证据,减少重复工作,提升效率。

AI Co-Scientist for Knowledge Synthesis in Medical Contexts: A Proof of Concept

  • 基于PICOS框架构建AI系统,实现研究要素自动提取
  • 模型对研究设计分类准确率达95.7%,比对专家标注一致性高
  • 适合需要高效整合文献的科研人员和政策制定者

生物医学研究中的资源浪费源于重复性研究、报告不全及传统证据合成流程难以扩展。本文提出一种基于显式形式化人口、干预、对照、结局与研究设计(PICOS)的AI协作者平台,集成关系型存储、向量语义检索与Neo4j知识图谱。在阿尔茨海默病-运动及非传染性疾病文献库上评估,使用Bi-LSTM基线与PubMedBERT微调的Transformer多任务分类器实现标题摘要中PICOS合规性与研究设计分类。全文合成采用混合向量与图检索的检索增强生成,结合BERTopic识别主题结构、冗余与证据空白。Transformer模型在研究设计分类上达到95.7%准确率,与专家标注高度一致;Bi-LSTM在PICOS合规检测中达87%准确率。检索增强生成在需结构约束、跨研究整合与图推理的任务中表现更优,而无检索生成在高层总结中仍具竞争力。主题建模揭示显著主题冗余,并识别出未充分研究领域。结果表明,具备PICOS感知与可解释性的自然语言处理可提升证据合成的可扩展性、透明度与效率。该架构具有领域通用性,为降低各生物医学领域的研究浪费提供实用框架。

原文摘要 · Abstract (English)

Research waste in biomedical science is driven by redundant studies, incomplete reporting, and the limited scalability of traditional evidence synthesis workflows. We present an AI co-scientist for scalable and transparent knowledge synthesis based on explicit formalization of Population, Intervention, Comparator, Outcome, and Study design (PICOS). The platform integrates relational storage, vector-based semantic retrieval, and a Neo4j knowledge graph. Evaluation was conducted on dementia-sport and non-communicable disease corpora. Automated PICOS compliance and study design classification from titles and abstracts were performed using a Bidirectional Long Short-Term Memory baseline and a transformer-based multi-task classifier fine-tuned from PubMedBERT. Full-text synthesis employed retrieval-augmented generation with hybrid vector and graph retrieval, while BERTopic was used to identify thematic structure, redundancy, and evidence gaps. The transformer model achieved 95.7% accuracy for study design classification with strong agreement against expert annotations, while the Bi-LSTM achieved 87% accuracy for PICOS compliance detection. Retrieval-augmented generation outperformed non-retrieval generation for queries requiring structured constraints, cross-study integration, and graph-based reasoning, whereas non-retrieval approaches remained competitive for high-level summaries. Topic modeling revealed substantial thematic redundancy and identified underexplored research areas. These results demonstrate that PICOS-aware and explainable natural language processing can improve the scalability, transparency, and efficiency of evidence synthesis. The proposed architecture is domain-agnostic and offers a practical framework for reducing research waste across biomedical disciplines.

AI科研证据合成医学NLP知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。