arXiv:2503.17784cs.AI2025-03AAAI被引 8

解决脑部CT报告生成中医学实体识别不均衡问题,提升报告准确性与全面性。

MEPNet: Medical Entity-balanced Prompting Network for Brain CT Report Generation

  • 引入医学实体视觉嵌入与学习状态评分,动态平衡各类解剖结构和病灶的建模
  • 在两个脑部CT报告生成数据集上,报告临床准确率与文本连贯性均显著提升
  • 适合关注医学影像报告生成、多模态模型优化的研究者与临床辅助系统开发者

自动生成脑部CT报告受到广泛关注,因其可辅助放射科医生诊断颅内疾病。然而,脑部CT包含大量医学实体(如不同解剖区域和病灶),其在三维体数据中的空间分布极不一致,导致现有方法对医学实体学习存在偏差,生成报告出现重复与不准确。为此,我们提出医学实体均衡提示网络(MEPNet),利用大语言模型(LLM)公平解析各类医学实体,实现精准的脑部CT报告生成。通过引入医学实体的视觉嵌入及其学习状态作为增强线索,引导LLM平衡不同实体的学习,从而生成更全面的报告。首先,设计知识驱动联合注意力机制,结合显式与隐式医学知识提取实体模式;其次,构建学习状态评分器,为每个实体生成独特学习状态评估;最后,将实体视觉嵌入与学习状态精巧融合进多模态提示,指导LLM自适应调整学习过程,覆盖更多详细发现。在两个脑部CT报告生成基准数据集上进行实验,验证了该方法在临床准确性和文本连贯性上的有效性。

原文摘要 · Abstract (English)

The automatic generation of brain CT reports has gained widespread attention, given its potential to assist radiologists in diagnosing cranial diseases. However, brain CT scans involve extensive medical entities, such as diverse anatomy regions and lesions, exhibiting highly inconsistent spatial patterns in 3D volumetric space. This leads to biased learning of medical entities in existing methods, resulting in repetitiveness and inaccuracy in generated reports. To this end, we propose a Medical Entity-balanced Prompting Network (MEPNet), which harnesses the large language model (LLM) to fairly interpret various entities for accurate brain CT report generation. By introducing the visual embedding and the learning status of medical entities as enriched clues, our method prompts the LLM to balance the learning of diverse entities, thereby enhancing reports with comprehensive findings. First, to extract visual embedding of entities, we propose Knowledge-driven Joint Attention to explore and distill entity patterns using both explicit and implicit medical knowledge. Then, a Learning Status Scorer is designed to evaluate the learning of entity visual embeddings, resulting in unique learning status for individual entities. Finally, these entity visual embeddings and status are elaborately integrated into multi-modal prompts, to guide the text generation of LLM. This process allows LLM to self-adapt the learning process for biased-fitted entities, thereby covering detailed findings in generated reports. We conduct experiments on two brain CT report generation benchmarks, showing the effectiveness in clinical accuracy and text coherence.

医学影像报告生成大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。