arXiv:2604.13328cs.LG2026-04

用轻量微调实现病理报告自动分期与标志物提取

Multi-Task LLM with LoRA Fine-Tuning for Automated Cancer Staging and Biomarker Extraction

  • 多任务并行分类头,结合LoRA微调提升效率
  • 宏F1达0.976,超越规则模型和单任务基线
  • 适合临床辅助决策与癌症研究数据构建

病理报告是乳腺癌分期的权威记录,但其非结构化格式阻碍了大规模数据整理。尽管大语言模型具备语义推理能力,但部署常受限于高计算成本和幻觉风险。本研究提出一种参数高效、多任务的自动化框架,用于提取肿瘤-淋巴结-转移(TNM)分期、组织学分级和生物标志物。我们在10,677份专家验证的报告上,使用低秩适应(LoRA)微调Llama-3-8B-Instruct编码器。不同于生成式方法,该架构采用并行分类头以确保模式一致性。实验表明,模型在宏F1得分上达到0.976,成功解决复杂上下文歧义与报告格式异质性问题,优于基于规则的自然语言处理流程、零样本大语言模型及单任务大语言模型基线。所提出的适配器高效、多任务架构,可实现可靠、可扩展的病理来源癌症分期与标志物分析,有望提升临床决策支持并加速数据驱动的肿瘤学研究。

原文摘要 · Abstract (English)

Pathology reports serve as the definitive record for breast cancer staging, yet their unstructured format impedes large-scale data curation. While Large Language Models (LLMs) offer semantic reasoning, their deployment is often limited by high computational costs and hallucination risks. This study introduces a parameter-efficient, multi-task framework for automating the extraction of Tumor-Node-Metastasis (TNM) staging, histologic grade, and biomarkers. We fine-tune a Llama-3-8B-Instruct encoder using Low-Rank Adaptation (LoRA) on a curated, expert-verified dataset of 10,677 reports. Unlike generative approaches, our architecture utilizes parallel classification heads to enforce consistent schema adherence. Experimental results demonstrate that the model achieves a Macro F1 score of 0.976, successfully resolving complex contextual ambiguities and heterogeneous reporting formats that challenge traditional extraction methods including rule-based natural language processing (NLP) pipelines, zero-shot LLMs, and single-task LLM baselines. The proposed adapter-efficient, multi-task architecture enables reliable, scalable pathology-derived cancer staging and biomarker profiling, with the potential to enhance clinical decision support and accelerate data-driven oncology research.

癌症分期多任务学习LoRA微调病理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。