arXiv:2608.14399cs.CYcs.AI2026-08

AI医生推荐受声誉影响大,性别种族隐性偏见难被模型察觉

Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice

  • 通过3024组虚拟医生卡片测试,量化声誉与人口属性对AI推荐的影响
  • 评分从3.9升至4.7,推荐概率提高31.4个百分点;费用从90涨至190美元,降低20.0个百分点
  • 女性及少数族裔名称获更高推荐,但模型自身解释中几乎不提此类因素

患者越来越多地使用大语言模型(LLM)助手选择医生,使这些系统成为隐形的AI信息中介:在人与人之间做出选择,悄然决定哪些医生被可见。我们报告了一项预先设定的随机算法审计,探究推动推荐结果的因果因素。七种模型(六种开源;gpt-4o-mini)在3024个选择集上从五张模拟全科医生卡片中做出选择,涉及三种患者角色、九种提示改写和九个实验组,共生成40,068条评分响应;姓名用于传递性别和种族信号,符合对照审计方法。声誉信号占主导:评分从3.9升至4.7,推荐概率提升31.4个百分点;费用从90美元增至190美元,下降20.0个百分点。尽管存在公平性问题,但偏差方向与人类研究预测相反:女性名称获得2.5个百分点优势,西班牙裔、南亚裔和非裔名称分别获得1.3至2.9个百分点优势,相当于每诊费多出7至14美元。首位无内容位置价值达11美元。然而,模型在最多0.03%的情况下提及性别或种族,且在0.39%的试验中完全回避,导致这些影响在自述理由中不可见。依赖模型自我报告的透明度要求无法捕捉这些偏差。一个推理模型直接未能通过预设可审计性标准。冻结设计确保审计可重复:任何新模型均可使用相同刺激进行评估,使行为审计而非自述解释,成为合适的监控技术。

原文摘要 · Abstract (English)

Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models' own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.

AI伦理医疗AI偏见检测模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。