arXiv:2410.01089cs.CV2024-10被引 19

首个评估医学多模态大模型公平性的基准,覆盖种族、性别等四类人群。

FMBench: Benchmarking Fairness in Multimodal Large Language Models on Medical Tasks

  • 构建包含四类人口属性的医疗多模态公平性评测基准。
  • 在零样本下评估8个模型,发现性能差异显著,最大偏差达27%。
  • 引入临床对齐评估指标,适合医疗AI研究者与政策制定者参考。

多模态大语言模型(MLLMs)在医学任务如视觉问答(VQA)和报告生成(RG)中表现优异,但其在不同人口群体间的公平性仍缺乏系统评估。这主要源于现有医学多模态数据集缺乏人口多样性。为此,我们提出FMBench,首个专门用于评估MLLMs在医疗任务中公平性的基准。该基准涵盖种族、民族、语言和性别四类人口属性,覆盖零样本下的VQA与RG任务;其中VQA为自由回答形式,提升真实场景适用性,并减少预设选项带来的偏见;采用词法指标与基于LLM的临床对齐指标,从语言准确性和临床合理性双重维度评估模型表现;同时引入公平性感知性能(FAP)指标量化模型在不同群体间的性能差异。我们全面评估了8个主流开源MLLMs(参数量7B至26B),包括通用与医疗专用模型。所有数据与代码将在论文接受后公开。

原文摘要 · Abstract (English)

Advancements in Multimodal Large Language Models (MLLMs) have significantly improved medical task performance, such as Visual Question Answering (VQA) and Report Generation (RG). However, the fairness of these models across diverse demographic groups remains underexplored, despite its importance in healthcare. This oversight is partly due to the lack of demographic diversity in existing medical multimodal datasets, which complicates the evaluation of fairness. In response, we propose FMBench, the first benchmark designed to evaluate the fairness of MLLMs performance across diverse demographic attributes. FMBench has the following key features: 1: It includes four demographic attributes: race, ethnicity, language, and gender, across two tasks, VQA and RG, under zero-shot settings. 2: Our VQA task is free-form, enhancing real-world applicability and mitigating the biases associated with predefined choices. 3: We utilize both lexical metrics and LLM-based metrics, aligned with clinical evaluations, to assess models not only for linguistic accuracy but also from a clinical perspective. Furthermore, we introduce a new metric, Fairness-Aware Performance (FAP), to evaluate how fairly MLLMs perform across various demographic attributes. We thoroughly evaluate the performance and fairness of eight state-of-the-art open-source MLLMs, including both general and medical MLLMs, ranging from 7B to 26B parameters on the proposed benchmark. We aim for FMBench to assist the research community in refining model evaluation and driving future advancements in the field. All data and code will be released upon acceptance.

多模态医疗AI公平性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。