arXiv:2512.10750cs.CV2025-12被引 1

用轻量微调让多模态大模型生成更准确的结肠镜报告。

LDP: Parameter-Efficient Fine-Tuning of Multimodal LLM for Medical Report Generation

  • 用LoRA实现轻量微调,大幅降低训练成本。
  • 在临床专家评分中达7.2/10,显著减少幻觉。
  • 适合医疗AI落地,支持真实场景部署。

结肠镜息肉诊断对早期结直肠癌发现至关重要,但传统自动化报告因高质量多模态医学数据稀缺而存在不一致和幻觉问题。为此,我们提出LDP框架,利用多模态大语言模型(MLLMs)生成专业息肉诊断报告。具体而言,我们构建了MMEndo数据集,包含专家标注的结肠镜图像-文本对。采用参数高效微调(LoRA)对Qwen2-VL-7B模型进行微调,并通过直接偏好优化(DPO)使其符合临床标准。大量实验表明,该方法在自动指标和严格的临床专家评估中均优于现有基线(医师评分达7.2/10),训练计算成本相比全量微调降低833倍。该方案为基层医疗提供了可扩展、临床可行的路径,且在IU-XRay数据集上进一步验证了其鲁棒性。

原文摘要 · Abstract (English)

Colonoscopic polyp diagnosis is pivotal for early colorectal cancer detection, yet traditional automated reporting suffers from inconsistencies and hallucinations due to the scarcity of high-quality multimodal medical data. To bridge this gap, we propose LDP, a novel framework leveraging multimodal large language models (MLLMs) for professional polyp diagnosis report generation. Specifically, we curate MMEndo, a multimodal endoscopic dataset comprising expert-annotated colonoscopy image-text pairs. We fine-tune the Qwen2-VL-7B backbone using Parameter-Efficient Fine-Tuning (LoRA) and align it with clinical standards via Direct Preference Optimization (DPO). Extensive experiments show that our LDP outperforms existing baselines on both automated metrics and rigorous clinical expert evaluations (achieving a Physician Score of 7.2/10), significantly reducing training computational costs by 833x compared to full fine-tuning. The proposed solution offers a scalable, clinically viable path for primary healthcare, with additional validation on the IU-XRay dataset confirming its robustness.

医疗AI多模态轻量微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。