arXiv:2511.01052cs.AIphysics.med-ph2025-11

用大模型从病理报告中自动提取癌症分期规则,无需大量标注数据。

Knowledge Elicitation with Large Language Models for Interpretable Cancer Stage Identification from Pathology Reports

  • 通过迭代提示和检索增强生成,让大模型从无标注报告中提炼分期规则。
  • 在乳腺癌数据上,两种方法均实现超过90%的分期准确率,优于传统模型。
  • 结果可解释性强,适合医疗场景中缺乏标注数据的临床应用。

癌症分期对患者预后和治疗方案制定至关重要,但从非结构化病理报告中提取病理学TNM分期仍具挑战。现有自然语言处理与机器学习方法通常依赖大规模标注数据,限制了其可扩展性与适应性。本文提出两种知识萃取方法:一是基于长期记忆的迭代提示(KEwLTM),直接从无标注报告中推导分期规则;二是基于检索增强生成的萃取(KEwRAG),预先从指南中提取规则并应用,提升可解释性且避免重复检索。我们利用大模型在预训练中获得的广泛知识,在TCGA乳腺癌病理报告数据集上评估其对T和N分期的识别性能,对比两种开源LLM上的多种基线方法。结果显示,当零样本思维链推理有效时,KEwLTM表现更优;反之,KEwRAG更具优势。两者均通过显式规则提供透明、可解释的接口。该研究证明知识萃取方法在低标注数据环境下具备高可扩展性与高性能,适用于临床场景。

原文摘要 · Abstract (English)

Cancer staging is critical for patient prognosis and treatment planning, yet extracting pathologic TNM staging from unstructured pathology reports poses a persistent challenge. Existing natural language processing (NLP) and machine learning (ML) strategies often depend on large annotated datasets, limiting their scalability and adaptability. In this study, we introduce two Knowledge Elicitation methods designed to overcome these limitations by enabling large language models (LLMs) to induce and apply domain-specific rules for cancer staging. The first, Knowledge Elicitation with Long-Term Memory (KEwLTM), uses an iterative prompting strategy to derive staging rules directly from unannotated pathology reports, without requiring ground-truth labels. The second, Knowledge Elicitation with Retrieval-Augmented Generation (KEwRAG), employs a variation of RAG where rules are pre-extracted from relevant guidelines in a single step and then applied, enhancing interpretability and avoiding repeated retrieval overhead. We leverage the ability of LLMs to apply broad knowledge learned during pre-training to new tasks. Using breast cancer pathology reports from the TCGA dataset, we evaluate their performance in identifying T and N stages, comparing them against various baseline approaches on two open-source LLMs. Our results indicate that KEwLTM outperforms KEwRAG when Zero-Shot Chain-of-Thought (ZSCOT) inference is effective, whereas KEwRAG achieves better performance when ZSCOT inference is less effective. Both methods offer transparent, interpretable interfaces by making the induced rules explicit. These findings highlight the promise of our Knowledge Elicitation methods as scalable, high-performing solutions for automated cancer staging with enhanced interpretability, particularly in clinical settings with limited annotated data.

癌症分期大模型可解释性病理报告

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。