arXiv:2410.14948cs.CLcs.CV2024-10被引 2

用半人工标注数据提升医疗多模态模型诊断能力

SemiHVision: Enhancing Medical Multimodal Models with a Semi-Human Annotated Dataset and Fine-Tuned Instruction Generation

  • 构建半人工标注数据集SemiHVision,融合人工与自动增强提升医学知识表达
  • 在SLAKE和VQA-RAD上超越公开模型(79.0%)和私有模型(55.7%)
  • 设计新评测基准JAMA临床挑战,更真实评估诊断推理能力

多模态大语言模型虽取得进展,但在医学领域受限于专业知识不足。现有模型在实验室表现良好,却难以应用于真实临床场景,研究与实践间存在显著差距。本文从数据收集、模型微调到评估全链路改进:提出SemiHVision数据集,结合人工标注与自动化增强技术,提升医学知识表示与诊断推理能力;基于2400小时H100 GPU训练的PMC-Cambrian-8B-AN,在SLAKE与VQA-RAD等传统基准上性能超越公开模型HuatuoGPT-Vision-34B(79.0% vs. 66.7%)及私有通用模型Claude3-Opus(55.7%)。在评估阶段发现传统基准无法反映真实临床能力,因此引入专为诊断推理设计的JAMA临床挑战基准。在此新基准上,PMC-Cambrian-AN达到GPT-4评分1.29,显著优于HuatuoGPT-Vision-34B(1.13)与Claude3-Opus(1.17),证明其卓越的临床推理能力。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have made significant strides, yet they face challenges in the medical domain due to limited specialized knowledge. While recent medical MLLMs demonstrate strong performance in lab settings, they often struggle in real-world applications, highlighting a substantial gap between research and practice. In this paper, we seek to address this gap at various stages of the end-to-end learning pipeline, including data collection, model fine-tuning, and evaluation. At the data collection stage, we introduce SemiHVision, a dataset that combines human annotations with automated augmentation techniques to improve both medical knowledge representation and diagnostic reasoning. For model fine-tuning, we trained PMC-Cambrian-8B-AN over 2400 H100 GPU hours, resulting in performance that surpasses public medical models like HuatuoGPT-Vision-34B (79.0% vs. 66.7%) and private general models like Claude3-Opus (55.7%) on traditional benchmarks such as SLAKE and VQA-RAD. In the evaluation phase, we observed that traditional benchmarks cannot accurately reflect realistic clinical task capabilities. To overcome this limitation and provide more targeted guidance for model evaluation, we introduce the JAMA Clinical Challenge, a novel benchmark specifically designed to evaluate diagnostic reasoning. On this benchmark, PMC-Cambrian-AN achieves state-of-the-art performance with a GPT-4 score of 1.29, significantly outperforming HuatuoGPT-Vision-34B (1.13) and Claude3-Opus (1.17), demonstrating its superior diagnostic reasoning abilities.

医疗AI多模态模型评测诊断推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。