用稀疏自编码器发现并干预大模型在医疗中对黑人患者的偏见。
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
- 通过分析模型潜空间,定位与黑人患者相关的敏感特征
- 激活该特征会显著增加模型预测患者
- 适合关注医疗AI公平性的研究者和从业者
大语言模型在医疗领域应用日益广泛,但可能加剧现有偏见。本文评估稀疏自编码器(SAEs)揭示并控制模型对患者种族与污名化概念之间关联的能力。在Gemma-2模型中,我们识别出与黑人群体相关的潜在特征,该特征在合理输入(如“非裔美国人”)和问题性词汇(如“监禁”)上均有激活。进一步实验表明,可利用此特征引导模型生成关于黑人患者的内容,并导致输出中出现不当关联,例如显著提高患者“具攻击性”的风险评分。我们测试了通过潜空间操控缓解偏见的可行性,结果表明在简单任务中有效,但在更复杂真实的临床场景中效果有限。总体而言,SAEs有助于识别医疗LLMs中的敏感依赖,但基于潜空间的偏见缓解在真实任务中作用有限。
原文摘要 · Abstract (English)
LLMs are increasingly being used in healthcare. This promises to free physicians from drudgery, enabling better care to be delivered at scale. But the use of LLMs in this space also brings risks; for example, such models may worsen existing biases. How can we spot when LLMs are (spuriously) relying on patient race to inform predictions? In this work we assess the degree to which Sparse Autoencoders (SAEs) can reveal (and control) associations the model has made between race and stigmatizing concepts. We first identify SAE latents in Gemma-2 models which appear to correlate with Black individuals. We find that this latent activates on reasonable input sequences (e.g., "African American") but also problematic words like "incarceration". We then show that we can use this latent to steer models to generate outputs about Black patients, and further that this can induce problematic associations in model outputs as a result. For example, activating the Black latent increases the risk assigned to the probability that a patient will become "belligerent". We evaluate the degree to which such steering via latents might be useful for mitigating bias. We find that this offers improvements in simple settings, but is less successful for more realistic and complex clinical tasks. Overall, our results suggest that: SAEs may offer a useful tool in clinical applications of LLMs to identify problematic reliance on demographics but mitigating bias via SAE steering appears to be of marginal utility for realistic tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。