将病理报告变成可检索的智能知识库,助力实时诊断决策。
PathoScribe: Transforming Pathology Data into a Living Library with a Unified LLM-Driven Framework for Semantic Retrieval and Clinical Integration
- 用统一大模型框架实现自然语言查询与推理,打通报告检索瓶颈。
- 在7万份报告上实现100%召回率,91.3%与人工一致的队列构建精度。
- 适合临床医生、研究者快速构建研究队列,大幅降低人力成本。
病理学是现代诊断和癌症治疗的基础,但其最宝贵资产——数百万份叙述性报告所承载的经验知识,仍难以获取。尽管机构正加速数字化病理流程,若缺乏有效的检索与推理机制,档案将沦为被动存储,无法真正支持患者诊疗。真正的进步不仅需要数字化,还需病理科医生在面对新诊断难题时,能实时查阅相似病例。本文提出 PathoScribe,一个统一的检索增强型大语言模型(LLM)框架,旨在将静态病理档案转化为可搜索、可推理的动态知识库。该系统支持自然语言病例探索、自动队列构建、临床问答、免疫组化(IHC)检测推荐及提示控制的报告转换。在7万份多机构外科病理报告上评估显示,自然语言检索的 Recall@10 达到100%,且检索驱动的推理质量优异(平均评分4.56/5)。关键的是,系统可从自由文本入选标准中自动构建研究队列,平均仅需9.2分钟,与人类评审者达成91.3%的一致性,且无有效病例被错误排除,相比传统手动审阅显著降低时间与成本。本工作为将数字病理档案转化为主动临床智能平台提供了可扩展基础。
原文摘要 · Abstract (English)
Pathology underpins modern diagnosis and cancer care, yet its most valuable asset, the accumulated experience encoded in millions of narrative reports, remains largely inaccessible. Although institutions are rapidly digitizing pathology workflows, storing data without effective mechanisms for retrieval and reasoning risks transforming archives into a passive data repository, where institutional knowledge exists but cannot meaningfully inform patient care. True progress requires not only digitization, but the ability for pathologists to interrogate prior similar cases in real time while evaluating a new diagnostic dilemma. We present PathoScribe, a unified retrieval-augmented large language model (LLM) framework designed to transform static pathology archives into a searchable, reasoning-enabled living library. PathoScribe enables natural language case exploration, automated cohort construction, clinical question answering, immunohistochemistry (IHC) panel recommendation, and prompt-controlled report transformation within a single architecture. Evaluated on 70,000 multi-institutional surgical pathology reports, PathoScribe achieved perfect Recall@10 for natural language case retrieval and demonstrated high-quality retrieval-grounded reasoning (mean reviewer score 4.56/5). Critically, the system operationalized automated cohort construction from free-text eligibility criteria, assembling research-ready cohorts in minutes (mean 9.2 minutes) with 91.3% agreement to human reviewers and no eligible cases incorrectly excluded, representing orders-of-magnitude reductions in time and cost compared to traditional manual chart review. This work establishes a scalable foundation for converting digital pathology archives from passive storage systems into active clinical intelligence platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。