构建儿科多模态问答基准,揭示大模型在儿童医疗中的年龄偏差问题
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
- 基于3417道文本题和2067道图像题构建儿科多模态评测集
- 大模型在婴幼儿组表现显著下降,年龄偏差严重
- 适合关注儿科AI公平性与医疗大模型评估的研究者
大型语言模型(LLMs)和视觉增强型语言模型(VLMs)在医学信息学、诊断与决策支持中取得显著进展。然而,这些模型存在系统性偏差,尤其是年龄偏差,导致其在面向儿童的文本与视觉问答任务中表现较差。这一偏差反映了医学研究中儿童领域投入不足的结构性失衡。为此,本文提出PediatricsMQA,一个涵盖3,417道文本多选题(覆盖131个儿科主题,七种发育阶段:孕前至青少年)和2,067道基于634张儿科图像的视觉多选题(来自67种影像模态、256个解剖区域)的综合性多模态儿科问答基准。数据通过人工与自动化结合的混合流程构建,整合了同行评审文献、已验证题库、现有基准及QA资源。对主流开源模型的评估显示,模型在年幼儿童群体中性能急剧下降,凸显了开发年龄感知方法以保障儿科医疗AI公平性的紧迫性。
原文摘要 · Abstract (English)
Large language models (LLMs) and vision-augmented LLMs (VLMs) have significantly advanced medical informatics, diagnostics, and decision support. However, these models exhibit systematic biases, particularly age bias, compromising their reliability and equity. This is evident in their poorer performance on pediatric-focused text and visual question-answering tasks. This bias reflects a broader imbalance in medical research, where pediatric studies receive less funding and representation despite the significant disease burden in children. To address these issues, a new comprehensive multi-modal pediatric question-answering benchmark, PediatricsMQA, has been introduced. It consists of 3,417 text-based multiple-choice questions (MCQs) covering 131 pediatric topics across seven developmental stages (prenatal to adolescent) and 2,067 vision-based MCQs using 634 pediatric images from 67 imaging modalities and 256 anatomical regions. The dataset was developed using a hybrid manual-automatic pipeline, incorporating peer-reviewed pediatric literature, validated question banks, existing benchmarks, and existing QA resources. Evaluating state-of-the-art open models, we find dramatic performance drops in younger cohorts, highlighting the need for age-aware methods to ensure equitable AI support in pediatric care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。