arXiv:2606.21194cs.CVcs.AI2026-06

首个医学多模态可解释性基准,评测AI能否用患者能懂的话描述影像结果。

MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models

论文配图:MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models
图 1 · 摘自论文原文
  • 构建分层语义框架,用医疗本体映射实现患者语言精准重构
  • 12万+图像样本验证:现有模型在通俗表达上严重退化
  • 轻量级评估器突破传统指标局限,贴近临床判断

医学视觉语言模型(Med-VLMs)在专家级表现优异,但其生成患者可理解描述的能力尚未充分探索。随着《21世纪治愈法案》要求立即向患者开放影像诊断结果,评估这些模型是否能弥合专家与患者之间的语言鸿沟,已成为关乎患者教育和共同决策的紧迫问题。为此,我们提出MedLayXPlain,首个大规模多模态基准与评估框架,用于医学通俗语言生成(MLLG)。MedLayXPlain-122K包含来自12个公开数据集的122,789个区域锚定样本,覆盖8种成像模态,每条样本配有专家与通俗描述,均基于三层次统一医学语言系统(UMLS)本体层级,涵盖7个语义组、43个语义类型和2,411个医学概念。通俗描述通过分层本体验证精炼(HOVER)流程生成,包括患者中心词汇映射、大模型约束重写及跨模型视觉验证,确保语义等价且避免幻觉。我们进一步提出MedLayEval,一个从270亿参数验证器蒸馏出的30亿参数轻量级评估器,从五个临床相关维度评分专家-通俗一致性,解决了标准自然语言生成指标与临床判断的相关性差的问题。在MedLayXPlain-122K上对33个视觉语言模型的评测显示,存在系统性专家-通俗差距:医学专用模型在专家描述上表现良好,但在通俗表达上显著退化;通用模型虽更易懂,却缺乏临床精确性,证实当前范式均无法有效支持面向患者的沟通。

原文摘要 · Abstract (English)

Medical Vision-Language Models (Med-VLMs) achieve strong expert-level performance, yet their ability to generate patient-accessible descriptions remains underexplored. With the 21st Century Cures Act now mandating immediate patient access to diagnostic imaging results, evaluating whether Med-VLMs can bridge this Expert-Lay Gap is both urgent and clinically consequential for patient education and shared decision-making. To this end, we introduce MedLayXPlain, the first large-scale multimodal benchmark and evaluation framework for Medical Lay Language Generation (MLLG). MedLayXPlain-122K provides 122,789 region-grounded samples across 8 imaging modalities from 12 publicly available source datasets, each comprising a medical image with paired expert and lay captions anchored in a three-level Unified Medical Language System (UMLS) ontology hierarchy spanning 7 semantic groups, 43 semantic types, and 2,411 medical concepts. Lay captions are constructed via Hierarchical Ontology-Verified Refinement (HOVER), a three-step pipeline combining patient-centric vocabulary mapping, LLM-based constrained rewriting, and cross-model visual verification to enforce semantic equivalence while preventing hallucination. We further introduce MedLayEval, a lightweight 3B evaluator distilled from a 27B verifier that scores expert-lay alignment across five clinically grounded attributes, addressing the poor correlation between standard NLG metrics and clinical judgment. Benchmarking 33 VLMs on MedLayXPlain-122K reveals a systematic Expert-Lay Gap: medical VLMs achieve strong expert captioning but suffer significant lay-register degradation, while general-purpose VLMs produce more accessible language yet lack clinical precision, confirming that neither current paradigm adequately serves patient-facing communication.

医学AI多模态可解释性患者沟通

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。