用多模态上下文学习生成病理图像报告,提升准确性和泛化能力。
Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning
- 基于训练集动态检索相似图像-报告对,增强上下文相关性。
- 在HistGen数据集上各项指标均达当前最优,覆盖多种疾病与报告长度。
- 适合医疗AI研究者与临床辅助系统开发者参考。
从病理图像自动生成医学报告是一项关键挑战,需要有效的视觉表征和领域知识。受人类专家实践启发,我们提出一种名为PathGenIC的上下文学习框架,将训练集中的上下文信息与多模态上下文学习(ICL)机制结合。该方法动态检索语义相似的全切片图像(WSI)-报告对,并引入自适应反馈以增强上下文相关性与生成质量。在HistGen基准上的评估显示,该框架在BLEU、METEOR和ROUGE-L等指标上均达到当前最优表现,且在不同报告长度和疾病类别下均表现出强鲁棒性。通过最大化训练数据利用效率,有效弥合视觉与语言模态的鸿沟,为人工智能驱动的病理报告生成提供了新方案,为未来多模态临床应用奠定坚实基础。
原文摘要 · Abstract (English)
Automating medical report generation from histopathology images is a critical challenge requiring effective visual representations and domain-specific knowledge. Inspired by the common practices of human experts, we propose an in-context learning framework called PathGenIC that integrates context derived from the training set with a multimodal in-context learning (ICL) mechanism. Our method dynamically retrieves semantically similar whole slide image (WSI)-report pairs and incorporates adaptive feedback to enhance contextual relevance and generation quality. Evaluated on the HistGen benchmark, the framework achieves state-of-the-art results, with significant improvements across BLEU, METEOR, and ROUGE-L metrics, and demonstrates robustness across diverse report lengths and disease categories. By maximizing training data utility and bridging vision and language with ICL, our work offers a solution for AI-driven histopathology reporting, setting a strong foundation for future advancements in multimodal clinical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。