构建真实中文在线问诊多模态数据集,评估大模型临床表现
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

- 从真实医患对话中提取临床关键场景,生成标准化多模态问答任务
- 5620个真实病例显示图像信息对临床判断至关重要,当前模型仍不及医生水平
- 专为医疗安全设计评分标准,适合评估医疗大模型的可靠性与安全性
大型语言模型在在线问诊中日益普及,但现有基准与真实临床实践脱节。多数依赖合成对话或患者模拟器,忽略患者上传的医学图像,且使用多选题或词重叠度评估开放问答,难以反映临床质量。我们提出MedRealMM,一个基于全国性中国互联网医院去标识化医患互动的大规模多模态基准。通过多模态临床挑战点(MCCP)提取框架,识别真实咨询轨迹中的临床难点,并转化为保留文本-图像上下文的下一轮回复生成任务。每例均配有医师修订的案例专用评分标准,奖励临床合理行为,惩罚不安全、无依据或矛盾的回答。当前版本包含覆盖64个临床科室的5620个真实多模态病例。我们评估了19个通用及医学专用大模型,包括纯文本和多模态系统。结果表明,图像信息对可靠临床表现至关重要,当前前沿模型仍低于在线医生表现。尽管部分模型满足的正面临床标准数量与医生相当甚至更多,但触发的负面标准也更多,说明安全敏感错误规避仍是核心瓶颈。MedRealMM为真实在线问诊中的多模态医疗推理提供了现实且可复现的评估基准。数据集将公开发布于Hugging Face:https://huggingface.co/datasets/jdh-algo/MedRealMM。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate open-ended clinical responses using multiple-choice or lexical-overlap metrics that poorly reflect clinical quality. We introduce \textbf{MedRealMM}, a large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collected from a nationwide Chinese internet hospital. MedRealMM uses a Multimodal Clinical Challenge Point (MCCP) extraction framework to identify clinically demanding moments in authentic consultation trajectories and converts each into a standardized next-response generation task while preserving the preceding text-image context. Each instance is paired with a case-specific rubric refined by physicians that rewards clinically desirable behaviors and penalizes unsafe, unsupported, or contradictory responses. The current release contains 5,620 real-world multimodal cases spanning 64 clinical departments. We evaluate 19 general-purpose and medical-specialized LLMs, including text-only and multimodal systems. Our results show that image information is critical for reliable clinical performance and that current frontier models remain below the online physician response. Although some frontier models satisfy as many or more positive clinical criteria than physicians, they trigger more negative criteria, indicating that safety-sensitive error avoidance remains a central bottleneck. MedRealMM offers a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultation. The dataset will be publicly available on Hugging Face at https://huggingface.co/datasets/jdh-algo/MedRealMM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。