构建急诊室视觉问答数据集,评估大模型在医院场景的实用能力
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
- 基于专家标注的开放问题,构建急诊室多场景图文问答数据集
- 多模型对比显示现有大模型在医疗场景下准确率仍不理想
- 适合医疗AI研究者、临床辅助系统开发者参考
全球医护人员短缺催生了智能医疗助手的需求,以在必要时协助监测与预警。本文通过在医院环境中开展视觉问答(VQA)任务,评估现有大视觉语言模型(LVLMs)的医疗知识水平。我们提出紧急病房视觉问答(ERVQA)数据集,包含涵盖多样化急诊场景的<图像, 问题, 答案>三元组,是首个面向LVLMs的急诊室基准数据集。通过构建详细错误分类体系并分析答案趋势,揭示该任务的复杂性。我们采用传统和改进的VQA评价指标(蕴含得分与CLIPScore置信度)对主流开源与闭源模型进行基准测试。通过对模型错误模式的分析,发现其表现受解码器类型、模型规模及上下文示例等属性影响。结果表明,ERVQA任务极具挑战性,凸显了开发专用领域解决方案的迫切需求。
原文摘要 · Abstract (English)
The global shortage of healthcare workers has demanded the development of smart healthcare assistants, which can help monitor and alert healthcare workers when necessary. We examine the healthcare knowledge of existing Large Vision Language Models (LVLMs) via the Visual Question Answering (VQA) task in hospital settings through expert annotated open-ended questions. We introduce the Emergency Room Visual Question Answering (ERVQA) dataset, consisting of <image, question, answer> triplets covering diverse emergency room scenarios, a seminal benchmark for LVLMs. By developing a detailed error taxonomy and analyzing answer trends, we reveal the nuanced nature of the task. We benchmark state-of-the-art open-source and closed LVLMs using traditional and adapted VQA metrics: Entailment Score and CLIPScore Confidence. Analyzing errors across models, we infer trends based on properties like decoder type, model size, and in-context examples. Our findings suggest the ERVQA dataset presents a highly complex task, highlighting the need for specialized, domain-specific solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。