arXiv:2507.14238cs.CLcs.AI2025-07被引 9

LLMs会因用户语言中的身份线索,给出不同待遇的医疗、薪资等建议。

Language Models Change Facts Based on the Way You Talk

  • 通过分析五类高风险应用,发现模型对用户性别、年龄、种族等有敏感反应
  • 同症状下不同族裔获得的医疗建议标准不同,年轻与年长者获政治事实答案不同
  • 推荐非白人更低薪资,女性比男性更高薪资,暴露系统性偏见

大型语言模型(LLMs)正被广泛应用于医疗咨询、求职建议等面向用户的场景。最新研究表明,这些模型能从极细微的语言模式中推断作者的身份信息。然而,关于模型如何在真实应用中利用这些身份线索仍知之甚少。本研究首次全面分析了身份标记在用户文本中如何影响五个高风险领域(医学、法律、政治、政府福利、工作薪资)的LLM响应。结果表明,模型对身份线索极为敏感,种族、性别和年龄持续影响输出。例如,在提供医疗建议时,相同症状下不同族裔获得的护理标准不同;针对政治事实问题,年长用户更可能得到保守立场的回答,年轻人则更可能得到自由立场的答案;在招聘场景中,非白人申请者被推荐更低薪资,女性则被推荐高于男性的薪资。这些偏差可能导致医疗不公、加剧薪酬差距,并为不同身份群体制造不同的事实认知。本文不仅揭示问题,还提供了评估语言中微妙身份编码对模型决策影响的新工具。鉴于严重后果,我们建议在部署前对所有面向用户的LLM应用进行类似全面评估。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being used in user-facing applications, from providing medical consultations to job interview advice. Recent research suggests that these models are becoming increasingly proficient at inferring identity information about the author of a piece of text from linguistic patterns as subtle as the choice of a few words. However, little is known about how LLMs use this information in their decision-making in real-world applications. We perform the first comprehensive analysis of how identity markers present in a user's writing bias LLM responses across five different high-stakes LLM applications in the domains of medicine, law, politics, government benefits, and job salaries. We find that LLMs are extremely sensitive to markers of identity in user queries and that race, gender, and age consistently influence LLM responses in these applications. For instance, when providing medical advice, we find that models apply different standards of care to individuals of different ethnicities for the same symptoms; we find that LLMs are more likely to alter answers to align with a conservative (liberal) political worldview when asked factual questions by older (younger) individuals; and that LLMs recommend lower salaries for non-White job applicants and higher salaries for women compared to men. Taken together, these biases mean that the use of off-the-shelf LLMs for these applications may cause harmful differences in medical care, foster wage gaps, and create different political factual realities for people of different identities. Beyond providing an analysis, we also provide new tools for evaluating how subtle encoding of identity in users' language choices impacts model decisions. Given the serious implications of these findings, we recommend that similar thorough assessments of LLM use in user-facing applications are conducted before future deployment.

大模型偏见身份识别医疗AI公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。