arXiv:2502.13319cs.CL2025-02EMNLP被引 10

用可解释性方法揭示医疗大模型中的性别和种族偏见

Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare

  • 通过激活分析定位模型中编码性别、种族信息的神经元
  • 干预性别相关激活可精准改变生成的临床描述与风险预测
  • 首次在医疗大模型中应用机制可解释性,适合研究公平性者参考

已有研究表明大语言模型会编码社会偏见,并在临床任务中体现。本文采用机制可解释性工具,揭示医疗场景下大模型中的社会人口学表征与偏见。具体问题为:能否识别模型中编码性别、种族等信息的激活?研究发现,性别信息高度集中于MLP层,可通过修复操作在推理时可靠操控,进而精准修改特定疾病的临床案例生成结果,并影响与性别相关的临床预测(如抑郁风险)。种族信息分布更广,但仍可部分干预。据我们所知,这是首个将机制可解释性方法应用于医疗大模型的研究。

原文摘要 · Abstract (English)

We know from prior work that LLMs encode social biases, and that this manifests in clinical tasks. In this work we adopt tools from mechanistic interpretability to unveil sociodemographic representations and biases within LLMs in the context of healthcare. Specifically, we ask: Can we identify activations within LLMs that encode sociodemographic information (e.g., gender, race)? We find that gender information is highly localized in MLP layers and can be reliably manipulated at inference time via patching. Such interventions can surgically alter generated clinical vignettes for specific conditions, and also influence downstream clinical predictions which correlate with gender, e.g., patient risk of depression. We find that representation of patient race is somewhat more distributed, but can also be intervened upon, to a degree. To our knowledge, this is the first application of mechanistic interpretability methods to LLMs for healthcare.

大模型偏见可解释性医疗AI性别偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。