MedTVT-R1用多模态大模型提升复杂疾病的诊断与推理能力。
MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
- 构建多模态融合框架,自适应加权不同医疗数据来源。
- 在多病诊断任务中表现优于现有方法,支持链式证据推理。
- 适合临床辅助诊断、报告生成及共病分析场景使用。
准确且可解释的多疾病诊断仍是医学研究中的关键挑战,尤其在处理异构多模态医疗数据时。当前方法多依赖单一模态数据,难以全面理解复杂疾病。为此,我们提出 MedTVT-R1,一种新型多模态大语言模型(MLLM)框架,旨在整合临床多模态数据以实现推理与多病诊断。我们构建了 MedTVT-QA 数据集,提供基于生理层面解释与疾病层面诊断的问答对,并采用链式证据(Chain of Evidence)方法。MedTVT-R1 引入模态感知层以捕捉模态间依赖关系,并自适应调整模态贡献权重。此外,我们采用基于组相对策略优化(GRPO)的强化学习微调,结合杰卡德奖励函数(Jaccard Reward)提升诊断推理能力。实验结果表明,MedTVT-R1 在多模态特征利用和多疾病诊断方面均表现出色,具有显著的临床应用潜力,如诊断报告生成与共病推理。代码与数据集已公开于 https://github.com/keke-nice/MedTVT-R1。
原文摘要 · Abstract (English)
Accurate and interpretable multi-disease diagnosis remains a critical challenge in medical research, particularly when leveraging heterogeneous multimodal medical data. Current approaches often rely on single-modal data, limiting their ability to comprehensively understand complex diseases. To address this, we propose MedTVT-R1, a novel Multimodal Large Language Model (MLLM) framework designed to integrate clinical multimodal data for reasoning and diagnosing multiple diseases. We construct MedTVT-QA, a curated instruction dataset that provides question-answer pairs for physiological-level interpretations and disease-level diagnoses with a Chain of Evidence approach. MedTVT-R1 incorporates a modality perception layer to capture inter-modal dependencies and adaptively weight modality contributions. Additionally, we employ Group Relative Policy Optimization (GRPO)-based Reinforcement Fine-Tuning with a Jaccard Reward function to enhance diagnostic reasoning. Experimental results demonstrate MedTVT-R1's superiority in multimodal feature utilization and multi-disease diagnosis, offering significant potential for clinical applications such as diagnostic report generation and comorbidity reasoning. The dataset and code are available at https://github.com/keke-nice/MedTVT-R1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。