arXiv:2604.05738cs.CL2026-04ACL被引 1

构建首个医学多模态通俗语义对齐基准,助力模型懂患者语言

MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models

论文配图:MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models
图 1 · 摘自论文原文
  • 基于概念锚定的精修流程,确保专业与通俗表述语义一致
  • 涵盖大规模多模态数据,支持临床专家与患者间的精准沟通评估
  • 适合医疗AI、人机交互研究者,推动可解释性医学应用

医学视觉语言模型(Med-VLMs)已在影像诊断理解上达到专家水平,但其训练主要依赖专业文献,难以用通俗语言传达结果以支持以患者为中心的照护。尽管文本简化研究已有进展,却缺乏大规模多模态基准来推动医学图像的通俗化理解。为此,我们提出MedLayBench-V——首个专注于专家-通俗语义对齐的大规模多模态基准。不同于易产生幻觉的简单简化方法,本数据集通过结构化概念锚定精修(SCGR)流程构建,结合统一医学语言系统(UMLS)的概念唯一标识符(CUIs)与微观实体约束,严格保证语义等价性。MedLayBench-V为下一代能弥合临床专家与患者沟通鸿沟的Med-VLMs提供了可验证的训练与评估基础。

原文摘要 · Abstract (English)

Medical Vision-Language Models (Med-VLMs) have achieved expert-level proficiency in interpreting diagnostic imaging. However, current models are predominantly trained on professional literature, limiting their ability to communicate findings in the lay register required for patient-centered care. While text-centric research has actively developed resources for simplifying medical jargon, there is a critical absence of large-scale multimodal benchmarks designed to facilitate lay-accessible medical image understanding. To bridge this resource gap, we introduce MedLayBench-V, the first large-scale multimodal benchmark dedicated to expert-lay semantic alignment. Unlike naive simplification approaches that risk hallucination, our dataset is constructed via a Structured Concept-Grounded Refinement (SCGR) pipeline. This method enforces strict semantic equivalence by integrating Unified Medical Language System (UMLS) Concept Unique Identifiers (CUIs) with micro-level entity constraints. MedLayBench-V provides a verified foundation for training and evaluating next-generation Med-VLMs capable of bridging the communication divide between clinical experts and patients.

医学多模态语义对齐可解释AI患者沟通

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。