arXiv:2502.19285cs.CV2025-02被引 1

清理病理报告文本可减少生成幻觉,提升报告准确性。

On the Importance of Text Preprocessing for Multimodal Representation Learning and Pathology Report Generation

  • 仅用显微镜可见特征描述训练模型,避免引入图像无法推断的信息。
  • 预处理后模型生成报告的幻觉显著减少,专家评估更可信。
  • 虽检索性能略降,但报告质量提升,适合临床报告生成场景。

病理学中的视觉-语言模型支持多模态病例检索与自动化报告生成。现有许多模型使用包含患者病史等图像无法推断信息的完整报告进行训练,可能导致生成报告出现幻觉。本文研究了从病理报告中筛选信息对多模态表征与报告生成质量的影响。我们对比了在完整报告和仅含基于H&E染色切片细胞/组织外观描述的预处理报告上训练的模型。实验基于BLIP-2框架,使用包含42,433张H&E全切片图像和19,636份对应病理报告的皮肤黑素细胞病变数据集。通过图像到文本、文本到图像检索以及专家病理医生的定性评估来衡量模型表现。结果表明,文本预处理能有效防止报告生成中的幻觉;尽管在跨模态检索性能上,完整报告训练模型表现更优,但预处理报告训练模型生成的报告质量更高。

原文摘要 · Abstract (English)

Vision-language models in pathology enable multimodal case retrieval and automated report generation. Many of the models developed so far, however, have been trained on pathology reports that include information which cannot be inferred from paired whole slide images (e.g., patient history), potentially leading to hallucinated sentences in generated reports. To this end, we investigate how the selection of information from pathology reports for vision-language modeling affects the quality of the multimodal representations and generated reports. More concretely, we compare a model trained on full reports against a model trained on preprocessed reports that only include sentences describing the cell and tissue appearances based on the H&E-stained slides. For the experiments, we built upon the BLIP-2 framework and used a cutaneous melanocytic lesion dataset of 42,433 H&E-stained whole slide images and 19,636 corresponding pathology reports. Model performance was assessed using image-to-text and text-to-image retrieval, as well as qualitative evaluation of the generated reports by an expert pathologist. Our results demonstrate that text preprocessing prevents hallucination in report generation. Despite the improvement in the quality of the generated reports, training the vision-language model on full reports showed better cross-modal retrieval performance.

病理报告多模态文本预处理幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。