arXiv:2505.17042cs.CLcs.CV2025-05被引 1

用图文模型生成放射科知识图谱,首次融合影像与报告信息。

VLM-KG: Multimodal Radiology Knowledge Graph Generation

  • 结合影像与报告的多模态框架,提升知识图谱生成效果。
  • 突破单模态局限,实现首个放射科多模态知识图谱生成方法。
  • 适合医学信息抽取、智能辅助诊断等临床研究者使用。

视觉-语言模型(VLMs)在自然语言生成方面表现优异,擅长指令遵循和结构化输出。知识图谱在放射科中至关重要,是事实信息的重要来源,并能提升多种下游任务性能。然而,由于放射科报告的专业语言及领域数据稀缺,生成专用于放射科的知识图谱面临巨大挑战。现有方法多为单模态,仅基于放射科报告生成知识图谱,忽略影像信息;且受限于上下文长度,难以处理长篇报告。为此,我们提出一种新型多模态VLM框架,用于放射科知识图谱生成。该方法超越已有方法,首次实现基于影像与文本的多模态放射科知识图谱生成。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated remarkable success in natural language generation, excelling at instruction following and structured output generation. Knowledge graphs play a crucial role in radiology, serving as valuable sources of factual information and enhancing various downstream tasks. However, generating radiology-specific knowledge graphs presents significant challenges due to the specialized language of radiology reports and the limited availability of domain-specific data. Existing solutions are predominantly unimodal, meaning they generate knowledge graphs only from radiology reports while excluding radiographic images. Additionally, they struggle with long-form radiology data due to limited context length. To address these limitations, we propose a novel multimodal VLM-based framework for knowledge graph generation in radiology. Our approach outperforms previous methods and introduces the first multimodal solution for radiology knowledge graph generation.

知识图谱多模态放射科VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。