LLM在医疗任务中存在显著公平性问题,不同人群表现差异大。
Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective
- 在六类医疗任务中测试主流LLM,评估其公平性表现。
- 不同人口群体间性能差异明显,且提供人口信息效果不一。
- 适合关注医疗AI公平性的研究者与从业者阅读。
本文研究大型语言模型(LLMs)在解决真实医疗任务时的表现,尤其从人口统计公平性角度出发。我们在六种不同的医疗任务上,评估了最先进的LLMs在三种主流学习框架下的表现,发现将LLMs应用于实际医疗场景存在显著挑战,且在不同人口群体间存在持续的公平性问题。研究还发现,显式提供人口信息的效果参差不齐,而模型自主推断此类信息则引发对偏见预测的担忧。即使让LLM作为具备实时指南访问权限的自主代理,也并不能保证性能提升。这些发现揭示了当前LLMs在医疗公平性方面的关键局限,凸显了该领域亟需专门研究。
原文摘要 · Abstract (English)
This paper studies the performance of large language models (LLMs), particularly regarding demographic fairness, in solving real-world healthcare tasks. We evaluate state-of-the-art LLMs with three prevalent learning frameworks across six diverse healthcare tasks and find significant challenges in applying LLMs to real-world healthcare tasks and persistent fairness issues across demographic groups. We also find that explicitly providing demographic information yields mixed results, while LLM's ability to infer such details raises concerns about biased health predictions. Utilizing LLMs as autonomous agents with access to up-to-date guidelines does not guarantee performance improvement. We believe these findings reveal the critical limitations of LLMs in healthcare fairness and the urgent need for specialized research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。