arXiv:2409.19676cs.CVcs.AI2024-09EMNLP被引 7

通过病灶线索增强视觉与文本对齐,提升脑部CT报告生成准确性

See Detail Say Clear: Towards Brain CT Report Generation via Pathological Clue-driven Representation Learning

  • 基于病灶区域、病变实体和报告主题构建病理线索,聚焦关键视觉信息
  • 在多个数据集上达到当前最优性能,报告生成质量显著提升
  • 适合作为临床辅助诊断系统的技术参考,尤其适合医学影像生成场景

脑部CT报告生成对颅内疾病诊断具有重要意义。现有研究致力于提升视觉与文本病理特征的一致性以改善报告连贯性,但仍面临两大挑战:1)冗余视觉表征——3D扫描中大量无关区域干扰模型对关键病灶的捕捉;2)语义表征偏移——有限的医学语料使模型难以将学习到的文本表征有效迁移到生成层。本文提出病灶线索驱动的表征学习(PCRL)模型,通过分割区域、病理实体和报告主题三个维度构建病理线索,全面捕捉视觉病理模式并学习跨模态特征表示。为适应文本生成任务,采用统一的大语言模型(LLM)并设计任务定制化指令,实现表征学习与报告生成间的无缝衔接。实验表明,该方法优于先前方法,在多个数据集上达到当前最优性能。代码已公开于https://github.com/Chauncey-Jheng/PCRL-MRG。

原文摘要 · Abstract (English)

Brain CT report generation is significant to aid physicians in diagnosing cranial diseases. Recent studies concentrate on handling the consistency between visual and textual pathological features to improve the coherence of report. However, there exist some challenges: 1) Redundant visual representing: Massive irrelevant areas in 3D scans distract models from representing salient visual contexts. 2) Shifted semantic representing: Limited medical corpus causes difficulties for models to transfer the learned textual representations to generative layers. This study introduces a Pathological Clue-driven Representation Learning (PCRL) model to build cross-modal representations based on pathological clues and naturally adapt them for accurate report generation. Specifically, we construct pathological clues from perspectives of segmented regions, pathological entities, and report themes, to fully grasp visual pathological patterns and learn cross-modal feature representations. To adapt the representations for the text generation task, we bridge the gap between representation learning and report generation by using a unified large language model (LLM) with task-tailored instructions. These crafted instructions enable the LLM to be flexibly fine-tuned across tasks and smoothly transfer the semantic representation for report generation. Experiments demonstrate that our method outperforms previous methods and achieves SoTA performance. Our code is available at "https://github.com/Chauncey-Jheng/PCRL-MRG".

脑部CT报告生成多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。