arXiv:2601.10945cs.CVcs.AI2026-01AAAI被引 1

用双模型对话模拟问诊,提升医学诊断准确率

PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis

  • 构建医生-患者双视觉语言模型,模拟多轮问诊过程
  • 合成症状经临床验证,与真实诊断高度匹配
  • 对话监督使模型诊断性能显著优于纯图像训练

传统医疗AI研究主要聚焦于图像分析,虽取得显著进展,但缺乏患者自述症状仍限制诊断准确性。为此,我们提出预问诊对话框架(PCDF),模拟真实诊疗流程:医生通过多轮提问逐步获取信息。具体而言,利用两个视觉语言模型(VLM)进行交互:DocVLM基于图像和对话历史生成追问问题,PatientVLM则根据真实诊断生成症状回复。我们对合成症状进行了小规模临床验证,专业医师确认其具备临床相关性、症状覆盖度和整体真实性。结果表明,该框架产生的图文对话与诊断高度一致,可用于微调DocVLM。相比仅依赖图像的训练方式,对话监督带来显著性能提升,凸显真实症状采集在诊断中的价值。

原文摘要 · Abstract (English)

Traditionally, AI research in medical diagnosis has largely centered on image analysis. While this has led to notable advancements, the absence of patient-reported symptoms continues to hinder diagnostic accuracy. To address this, we propose a Pre-Consultation Dialogue Framework (PCDF) that mimics real-world diagnostic procedures, where doctors iteratively query patients before reaching a conclusion. Specifically, we simulate diagnostic dialogues between two vision-language models (VLMs): a DocVLM, which generates follow-up questions based on the image and dialogue history, and a PatientVLM, which responds using a symptom profile derived from the ground-truth diagnosis. We additionally conducted a small-scale clinical validation of the synthetic symptoms generated by our framework, with licensed clinicians confirming their clinical relevance, symptom coverage, and overall realism. These findings indicate that the resulting DocVLM-PatientVLM interactions form coherent, multi-turn consultations paired with images and diagnoses, which we then use to fine-tune the DocVLM. This dialogue-based supervision leads to substantial gains over image-only training, highlighting the value of realistic symptom elicitation for diagnosis.

医学诊断对话系统视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。