arXiv:2608.23143cs.CV2026-08

用自研模型直接生成德语前列腺病理报告,解决多语言适配难题

An end-to-end-trained vision-language model for native-language prostate pathology report generation

论文配图:An end-to-end-trained vision-language model for native-language prostate pathology report generation
图 1 · 摘自论文原文
  • 从零训练多语言模型,不依赖英语编码器
  • 17,344对图像文本数据自动生成,恶性肿瘤识别F1达96.2%
  • 支持本地病历库训练,适合非英语医疗场景

前列腺癌是全球最常见的恶性肿瘤之一,每例活检核心的结构化报告给病理科医生带来负担。现有工具将其视为分类任务,需医生自行整合成完整报告;而多数基于图像-语言的模型依赖英语编码器,在其他临床语言中表现不佳。本文提出一种滑片级、端到端训练的视觉-语言模型,从构建上即具备语言独立性:分词器与模型均从零训练,本研究以德语为例进行演示。为缓解成对数据稀缺问题,设计自动化流程,利用本地部署的大语言模型将复合报告拆分为针对每个核心的图像-文本对,共生成17,344对数据,来自2,402例历史病例,无需人工标注。评估聚焦临床属性而非语言相似性,模型在恶性肿瘤检测上达到96.2% F1,在格里森分级上达65.2%,性能可媲美经FDA认证的分类器。分级结果在三个外部队列中通过潜在空间增强进一步验证。机构可基于自有病历库训练原生语言报告模型。

原文摘要 · Abstract (English)

Prostate cancer is among the most frequently diagnosed malignancies worldwide, and structured reporting of each biopsy core burdens pathologists. Existing tools frame this as classification, leaving pathologists to assemble coherent reports, while many slide-level vision-language models rely on English-centric encoders that transfer poorly to other clinical languages. We present a slide-level framework generating prostate biopsy reports that is language-independent by construction: tokenizer and model are trained from scratch, demonstrated here in German. To address paired-data scarcity, an automated pipeline uses a locally deployed large language model to split composite reports into core-specific image-text pairs, yielding 17,344 pairs from 2,402 historical cases without manual annotation. Evaluated for clinical attributes rather than linguistic similarity, the model achieves 96.2% F1 for malignancy detection and 65.2% for Gleason grading, competitive with an FDA-cleared classifier. Grading is further validated on three external cohorts with latent-space augmentation. Institutions can thus train native-language reporting models on their own archives.

病理报告生成多语言模型视觉语言模型前列腺癌

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。