专为肿瘤学设计轻量高效语言模型,支持跨语言迁移与多模态数据处理。
Towards Scalable and Cross-Lingual Specialist Language Models for Oncology
- 融合指令微调、检索增强生成与图谱知识,构建肿瘤领域专用NLP框架。
- 在实体识别、分期标注和病理报告分类任务中表现优异,准确率超基线模型。
- 仅用少量德语数据即可实现跨语言知识迁移,适合资源有限的医疗场景。
临床肿瘤学产生大量非结构化数据,常含不一致、缺失和模糊信息,难以支持数据驱动决策。通用大语言模型因缺乏领域特定推理能力,难以应对专业术语、上下文依赖解释及多模态数据融合等挑战。本文提出一种面向肿瘤学的轻量化、可适应的NLP框架,结合指令微调、检索增强生成(RAG)与基于图的知识集成。该框架在肿瘤特定任务中表现良好,包括命名实体识别(如癌症诊断识别)、实体链接(如关联标准本体)、TNM分期、文档分类(如从病理报告中识别癌种亚型)及治疗反应预测。模型强调可适应性与资源效率,仅使用来自苏黎世大学医院(USZ)的少量德语指令数据,验证了非英语数据对跨语言知识迁移的有效性。实验表明,在多个肿瘤学数据集上,模型在实体识别、关系抽取与文档分类任务中均取得优异表现。
原文摘要 · Abstract (English)
Clinical oncology generates vast, unstructured data that often contain inconsistencies, missing information, and ambiguities, making it difficult to extract reliable insights for data-driven decision-making. General-purpose large language models (LLMs) struggle with these challenges due to their lack of domain-specific reasoning, including specialized clinical terminology, context-dependent interpretations, and multi-modal data integration. We address these issues with an oncology-specialized, efficient, and adaptable NLP framework that combines instruction tuning, retrieval-augmented generation (RAG), and graph-based knowledge integration. Our lightweight models prove effective at oncology-specific tasks, such as named entity recognition (e.g., identifying cancer diagnoses), entity linking (e.g., linking entities to standardized ontologies), TNM staging, document classification (e.g., cancer subtype classification from pathology reports), and treatment response prediction. Our framework emphasizes adaptability and resource efficiency. We include minimal German instructions, collected at the University Hospital Zurich (USZ), to test whether small amounts of non-English language data can effectively transfer knowledge across languages. This approach mirrors our motivation for lightweight models, which balance strong performance with reduced computational costs, making them suitable for resource-limited healthcare settings. We validated our models on oncology datasets, demonstrating strong results in named entity recognition, relation extraction, and document classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。