arXiv:2503.06894cs.CVcs.AI2025-03中稿 · International Conf…

用视觉变压器+GPT-2生成病理图像描述,辅助医生发现细微病变。

A Deep Learning Approach for Augmenting Perceptional Understanding of Histopathology Images

  • 融合ViT与GPT-2的多模态模型,生成精准临床语境下的图像描述。
  • 在ARCH数据集上微调,能捕捉组织形态、染色差异等复杂病理特征。
  • 适合医学影像分析、辅助诊断系统开发人员参考使用。

近年来,数字技术在提升人类健康认知与感知能力方面取得显著进展,尤其在计算病理学领域。本文提出一种新方法,通过结合视觉变压器(Vision Transformers, ViT)与GPT-2的多模态模型,实现病理图像的图像描述生成。该模型在专用的ARCH数据集上进行微调,该数据集包含从临床与学术资源中提取的密集图像描述,涵盖组织形态、染色差异及病理状态等复杂特征。通过生成准确且具有上下文意义的描述,模型增强了医疗专业人员的认知能力,提升了疾病分类、分割与检测效率。该方法有助于发现原本可能被忽略的细微病理特征,从而提高诊断准确性。本研究展示了数字技术在增强人类医学图像分析认知能力方面的潜力,为实现更个性化、精准的医疗结果提供了支持。

原文摘要 · Abstract (English)

In Recent Years, Digital Technologies Have Made Significant Strides In Augmenting-Human-Health, Cognition, And Perception, Particularly Within The Field Of Computational-Pathology. This Paper Presents A Novel Approach To Enhancing The Analysis Of Histopathology Images By Leveraging A Mult-modal-Model That Combines Vision Transformers (Vit) With Gpt-2 For Image Captioning. The Model Is Fine-Tuned On The Specialized Arch-Dataset, Which Includes Dense Image Captions Derived From Clinical And Academic Resources, To Capture The Complexities Of Pathology Images Such As Tissue Morphologies, Staining Variations, And Pathological Conditions. By Generating Accurate, Contextually Captions, The Model Augments The Cognitive Capabilities Of Healthcare Professionals, Enabling More Efficient Disease Classification, Segmentation, And Detection. The Model Enhances The Perception Of Subtle Pathological Features In Images That Might Otherwise Go Unnoticed, Thereby Improving Diagnostic Accuracy. Our Approach Demonstrates The Potential For Digital Technologies To Augment Human Cognitive Abilities In Medical Image Analysis, Providing Steps Toward More Personalized And Accurate Healthcare Outcomes.

病理图像多模态视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。