临床大模型决策的公平性审计需设每步动作的不稳定性基准线
Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor
- 通过重复测试发现同一患者描述下模型动作仍会变动,揭示内在不稳定性
- 平均8.7%的动作变化率构成基准线,某些操作最高达17.9%的波动
- 建议所有公平性评估必须配比该基准线,否则结果不可信
反事实审计是检验临床智能体是否对不同人口特征但临床状况相同的患者区别对待的标准工具,其报告翻转率:仅患者描述改变时行动是否变化。我们发现该数值单独无意义。在十六个病例中重复相同条件十次(相同叙述、相同描述字符串,无变量),临床智能体动作在8.7%的结局-案例单元中发生变化,且不稳定性在各动作间差异显著,达八倍之差,从重症监护升级的0.022到管制药物警示的0.179。数据中任何人口学对比均无法区分于该基线。另一模型给出6.7%的合并基线,并近乎一致排序六个动作(斯皮尔曼相关0.94,精确p=0.017),表明该基线非单一系统偏差。五次采样多数投票可消除39%的不稳定性,之后趋于平缓;零模型模拟显示残余源于各单元异质率,故复现可缓解但无法根除。因此,未附每动作基线的反事实公平性估计均无法作为差异证据。测量基于FairMedAgent——一个评估临床大模型行为偏倚的评测框架,其估算量为范围内反事实翻转率,仅计录发布决策规则允许与临床医生裁定之间的翻转。该估算需带范围裁定,目前正在进行中;本文不宣称任何偏倚结论。每个合成病例执行六阶段轨迹(五个模型决策围绕确定性环境步),固定形式条件涵盖种族、性别、年龄、保险、英语能力及其交集。框架、基线协议及所有分析脚本均已开源。
原文摘要 · Abstract (English)
Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip rate: how often an action changes when only the patient descriptor changes. We show that this quantity is uninterpretable on its own. Re-running an identical condition ten times over sixteen vignettes (same narrative, same descriptor string, nothing varied) moved a clinical agent's action in 8.7% of outcome-vignette cells, and instability was heterogeneous across actions by a factor of eight, from 0.022 for ICU escalation to 0.179 for controlled-substance caution. No demographic contrast in our data was distinguishable from that floor. A second model gives a pooled floor of 6.7% and ranks the six actions almost identically (Spearman 0.94, exact p=0.017), so the floor is not one system's artefact. Majority-vote aggregation over five draws removes 39% of it and then flattens, and a null simulation attributes the residue to heterogeneous per-cell rates, so replication mitigates without eliminating. Any counterfactual fairness estimate reported without a per-action floor beside it therefore cannot be read as evidence of disparity. The measurements were taken with FairMedAgent, an evaluation harness for disparity in the actions of clinical LLM agents whose estimand, the within-range counterfactual flip rate, counts only flips between actions a published decision rule admits and a clinician has adjudicated. That estimand requires band adjudication, which is under way; no disparity result is claimed here. Each synthetic vignette runs a six-stage trajectory (five model-facing decisions around a deterministic environment step) under fixed-form conditions spanning race, sex, age, insurance, English proficiency, and their intersections. The harness, the floor protocol, and every analysis script are released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。