AI辅助病理医生更准判断免疫组化染色结果
iSight: Towards expert-AI co-assessment for improved immunohistochemistry staining interpretation
- 用全切片图像与组织元数据融合,多任务学习预测染色特征
- 在染色位置/强度/数量上准确率超85%/76%/75%,优于现有模型
- 专家与AI协同评估显著提升诊断一致性和准确性
免疫组化(IHC)用于分析组织切片中蛋白表达,常用于病理诊断和疾病分层。尽管针对H&E染色切片的AI模型已有进展,但其在IHC上的应用受限于领域特异性差异。本文提出HPA10M数据集,包含来自人类蛋白图谱的10,495,672张IHC图像,涵盖45种正常组织和20种主要癌种,并构建了iSight多任务学习框架,通过令牌级注意力机制融合全切片图像视觉特征与组织元数据,同时预测染色强度、位置、数量、组织类型及恶性程度。在独立测试数据上,iSight在位置、强度、数量预测上的准确率分别为85.5%、76.6%、75.7%,优于微调的基线模型(PLIP、CONCH)2.5–10.2%。其预测校准性良好,期望校准误差为0.0150–0.0408。八名病理科医生对两个数据集共200张图像进行评估,iSight在独立数据集上的表现优于初始医生判断(位置:79% vs 68%,强度:70% vs 57%,数量:68% vs 52%)。专家间一致性也提升,HPA数据集的科恩κ从0.63升至0.70,Stanford TMAD数据集从0.74升至0.76,表明专家- AI协同评估可有效改善IHC解读。本研究为提升IHC诊断准确性的AI系统奠定基础,展示了iSight融入临床流程以增强评估一致性与可靠性的潜力。
原文摘要 · Abstract (English)
Immunohistochemistry (IHC) provides information on protein expression in tissue sections and is commonly used to support pathology diagnosis and disease triage. While AI models for H\&E-stained slides show promise, their applicability to IHC is limited due to domain-specific variations. Here we introduce HPA10M, a dataset that contains 10,495,672 IHC images from the Human Protein Atlas with comprehensive metadata included, and encompasses 45 normal tissue types and 20 major cancer types. Based on HPA10M, we trained iSight, a multi-task learning framework for automated IHC staining assessment. iSight combines visual features from whole-slide images with tissue metadata through a token-level attention mechanism, simultaneously predicting staining intensity, location, quantity, tissue type, and malignancy status. On held-out data, iSight achieved 85.5\% accuracy for location, 76.6\% for intensity, and 75.7\% for quantity, outperforming fine-tuned foundation models (PLIP, CONCH) by 2.5--10.2\%. In addition, iSight demonstrates well-calibrated predictions with expected calibration errors of 0.0150-0.0408. Furthermore, in a user study with eight pathologists evaluating 200 images from two datasets, iSight outperformed initial pathologist assessments on the held-out HPA dataset (79\% vs 68\% for location, 70\% vs 57\% for intensity, 68\% vs 52\% for quantity). Inter-pathologist agreement also improved after AI assistance in both held-out HPA (Cohen's $κ$ increased from 0.63 to 0.70) and Stanford TMAD datasets (from 0.74 to 0.76), suggesting expert--AI co-assessment can improve IHC interpretation. This work establishes a foundation for AI systems that can improve IHC diagnostic accuracy and highlights the potential for integrating iSight into clinical workflows to enhance the consistency and reliability of IHC assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。