用图约束推理链生成病理报告,提升可解释性与准确性
Decomposing Whole Slide Image Report Generation with Graph-Constrained Multiple Instance Learning Workflows

- 将病理图像分块后通过多实例学习分类,逐个回答器官特异性问题
- 引入器官相关图结构约束答案组合,使报告生成链式推理更合理,链式Jaccard达0.702
- 适合关注医学影像可解释性、病理报告生成的临床与研究者
全切片图像(WSI)报告生成需识别空间分布的病理特征并组织成连贯诊断叙述。尽管直接的视觉-文本模型能生成流畅报告,但掩盖了视觉识别、结构化推理和语言生成各环节的贡献与失败模式。本文提出一种分解式框架:使用冻结的Virchow2图像块嵌入,通过多实例学习(MIL)分类头回答器官特异性诊断问题;再以器官条件化的图结构约束这些答案的组装过程,形成结构化推理链,最终由语言模型生成病理报告。在REG2026独立测试集(2,028张切片)上,该流程链式Jaccard得分为0.702。若移除图结构,性能降至0.420;用单一无器官区分的图则为0.398;若语言模型自由组合MIL预测结果,仅0.371。相同报告生成器下,图结构推理使报告得分从0.330提升至0.495。在350张外部TCGA WSI(涵盖7个REG器官)上,正确器官图在64.0%情况下被选中,前三位包含率86.6%。使用正确图结构使与粗粒度TCGA主诊断标签的一致性从61.8%提升至92.6%,表明器官路由是领域迁移下的主要瓶颈。整体而言,器官条件化的图约束推理链显著提升了结构化推理能力与报告质量,并支持特定阶段的错误定位。
原文摘要 · Abstract (English)
Whole-slide image (WSI) report generation requires recognizing spatially distributed pathological features and organizing them into a coherent diagnostic narrative. Although direct vision-to-text models can yield fluent reports, they obscure the contributions and failure modes of visual recognition, structured reasoning, and language generation. We propose a decomposed framework in which frozen Virchow2 tile embeddings are aggregated by multiple-instance learning (MIL) classification heads that answer organ-specific diagnostic questions. An organ-conditioned graph constrains the assembly of these answers into a structured reasoning chain, which a language model realizes as a pathology report. On the REG2026 held-out set of 2,028 slides, the proposed workflow achieved a chain-Jaccard score of 0.702. Performance fell to 0.420 without graph-based chain construction, 0.398 when the organ-specific graphs were replaced by a single organ-agnostic graph, and 0.371 when the language model constructed the chain freely from MIL predictions. Using the same report generator, graph-structured chains improved the report score from 0.330 to 0.495. On 350 external TCGA WSIs spanning the seven REG organs without fine-tuning, the expected organ graph was selected in 64.0% of cases and ranked among the top three in 86.6%. Providing the correct organ graph increased agreement with coarse TCGA primary-diagnosis labels from 61.8% to 92.6%, identifying organ routing as a main bottleneck under domain shift. Overall, organ-conditioned, graph-constrained chain assembly improves structured reasoning and report generation while enabling stage-specific error localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。