用健康社会决定因素分析大模型性别偏见,发现模型依赖刻板印象做判断。
Investigating Gender Stereotypes in Large Language Models via Social Determinants of Health
- 通过健康社会因素输入探测模型性别偏见
- 模型在诊断时依赖嵌入的性别刻板印象
- 适合关注医疗AI公平性的研究者阅读
大型语言模型(LLMs)在自然语言处理任务中表现优异,但常传播训练数据中的偏见,这在医疗等敏感领域影响深远。现有基准多聚焦单一社会决定因素(如性别或种族),忽视各因素间的交互作用,且缺乏情境化评估。本研究通过法语患者病历,探究性别与其他健康社会决定因素(SDoH)的相互关系。实验表明,可通过引入SDoH信息探测模型中的刻板印象;模型在决策中确实依赖这些嵌入的性别刻板印象。结果提示,评估多个SDoH因素的交互作用,可有效补充现有模型性能与偏见评估方法。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel in Natural Language Processing (NLP) tasks, but they often propagate biases embedded in their training data, which is potentially impactful in sensitive domains like healthcare. While existing benchmarks evaluate biases related to individual social determinants of health (SDoH) such as gender or ethnicity, they often overlook interactions between these factors and lack context-specific assessments. This study investigates bias in LLMs by probing the relationships between gender and other SDoH in French patient records. Through a series of experiments, we found that embedded stereotypes can be probed using SDoH input and that LLMs rely on embedded stereotypes to make gendered decisions, suggesting that evaluating interactions among SDoH factors could usefully complement existing approaches to assessing LLM performance and bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。