arXiv:2410.14763cs.CLcs.AI2024-10被引 8

自动生成医学偏见测试用例,提升医疗大模型评估规模与准确性

Enabling Scalable Evaluation of Bias Patterns in Medical LLMs

  • 基于医学知识图谱自动构建偏见测试用例
  • 生成的测试用例能有效揭示医疗大模型中的偏见模式
  • 适合医疗AI安全评估与负责任部署的研究者使用

大型语言模型(LLMs)在应对多种医疗挑战方面展现出巨大潜力。然而,在医疗等高风险场景中部署时,其潜在的偏见行为可能引发对个体的不公平对待。为实现医疗大模型的负责任应用,严格评估至关重要。由于医疗场景的复杂性和多样性,现有研究主要依赖人工构建的数据集进行偏见评估。本文提出一种新方法,通过基于严谨医学证据的自动化测试用例生成,实现偏见评估的规模化。针对偏见表征的领域特异性、生成过程中的幻觉问题以及健康结果与敏感属性间的多重依赖关系,我们整合医学知识图谱、医学本体及定制化通用大模型评估框架,构建生成流水线。大量实验证明,该方法生成的测试用例可在更大范围和更灵活条件下揭示医疗大模型的偏见模式。我们发布了基于该流水线构建的大规模偏见评估数据集,涵盖若干医学案例研究。应用演示可通过 https://vignette.streamlit.app 查看,代码开源地址为 https://github.com/healthylaife/autofair。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown impressive potential in helping with numerous medical challenges. Deploying LLMs in high-stakes applications such as medicine, however, brings in many concerns. One major area of concern relates to biased behaviors of LLMs in medical applications, leading to unfair treatment of individuals. To pave the way for the responsible and impactful deployment of Med LLMs, rigorous evaluation is a key prerequisite. Due to the huge complexity and variability of different medical scenarios, existing work in this domain has primarily relied on using manually crafted datasets for bias evaluation. In this study, we present a new method to scale up such bias evaluations by automatically generating test cases based on rigorous medical evidence. We specifically target the challenges of a) domain-specificity of bias characterization, b) hallucinating while generating the test cases, and c) various dependencies between the health outcomes and sensitive attributes. To that end, we offer new methods to address these challenges integrated with our generative pipeline, using medical knowledge graphs, medical ontologies, and customized general LLM evaluation frameworks in our method. Through a series of extensive experiments, we show that the test cases generated by our proposed method can effectively reveal bias patterns in Med LLMs at larger and more flexible scales than human-crafted datasets. We publish a large bias evaluation dataset using our pipeline, which is dedicated to a few medical case studies. A live demo of our application for vignette generation is available at https://vignette.streamlit.app. Our code is also available at https://github.com/healthylaife/autofair.

医疗AI偏见评估自动化测试大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。