arXiv:2604.14656cs.AIcs.CL2026-04

构建多轮多模态医学教育评测基准,让AI像医生一样图文并茂地解释影像报告。

Rethinking Patient Education as Multi-turn Multi-modal Interaction

论文配图:Rethinking Patient Education as Multi-turn Multi-modal Interaction
图 1 · 摘自论文原文
  • 设计多轮交互流程,结合图像与语言生成可解释的患者教育内容。
  • 发现现有模型在视觉定位上常偏离报告证据,安全性和情绪应对能力最弱。
  • 适合研究医疗AI对话、多模态生成与健康素养提升的学者使用。

当前多数医学多模态评测聚焦静态任务如图像问答、报告生成和通俗化改写。但患者教育更具挑战性:系统需从影像中识别相关证据,引导患者关注位置,用易懂语言解释发现,并应对困惑或焦虑。然而多数研究仍仅限文本,尽管图文结合更利于理解。本文提出MedImageEdu,一个面向多轮、基于证据的放射科患者教育评测基准。每例包含影像报告(文本+图像),由DoctorAgent与受隐藏个人特征(如教育水平、健康素养、性格)影响的PatientAgent交互。当问题需视觉支持时,DoctorAgent可向内置绘图工具发出基于报告、图像和当前问题的绘制指令,工具返回图像后,DoctorAgent生成含图像与口语化解释的最终多模态响应。该基准包含150例,来自三个数据源,从五个维度评估:咨询过程、安全性与覆盖范围、语言质量、绘图质量、图文响应质量。在代表性开闭源视觉语言模型代理中,我们发现三类共性差距:流畅语言常超越忠实视觉定位,安全性在各类疾病中均为最弱环节,情绪紧张交互比低教育或低健康素养更难处理。MedImageEdu为评估多模态代理是否真正基于证据教学提供了可控测试平台。

原文摘要 · Abstract (English)

Most medical multimodal benchmarks focus on static tasks such as image question answering, report generation, and plain-language rewriting. Patient education is more demanding: systems must identify relevant evidence across images, show patients where to look, explain findings in accessible language, and handle confusion or distress. Yet most patient education work remains text-only, even though combined image-and-text explanations may better support understanding. We introduce MedImageEdu, a benchmark for multi-turn, evidence-grounded radiology patient education. Each case provides a radiology report with report text and case images. A DoctorAgent interacts with a PatientAgent, conditioned on a hidden profile that captures factors such as education level, health literacy, and personality. When a patient question would benefit from visual support, the DoctorAgent can issue drawing instructions grounded in the report, case images, and the current question to a benchmark-provided drawing tool. The tool returns image(s), after which the DoctorAgent produces a final multimodal response consisting of the image(s) and a grounded plain-language explanation. MedImageEdu contains 150 cases from three sources and evaluates both the consultation process and the final multimodal response along five dimensions: Consultation, Safety and Scope, Language Quality, Drawing Quality, and Image-Text Response Quality. Across representative open- and closed-source vision-language model agents, we find three consistent gaps: fluent language often outpaces faithful visual grounding, safety is the weakest dimension across disease categories, and emotionally tense interactions are harder than low education or low health literacy. MedImageEdu provides a controlled testbed for assessing whether multimodal agents can teach from evidence rather than merely answer from text.

多模态患者教育医学AI交互式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。