用大模型模拟医疗决策,发现部分模型存在偏见
Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare
- 通过人口特征提示工程构建数字孪生受访者
- Llama 3 更准确反映种族与收入差异,但引入新偏见
- 提醒研究者注意模型与提示策略带来的双重偏差
生成式代理正被广泛用于模拟人类行为,依托大语言模型(LLMs)。这些数字代理为研究人类行为提供安全沙盒,不涉及隐私风险。本研究将理解美国研究(UAS)中关于医疗决策的调查数据,与生成式代理的模拟回答进行对比。通过基于人口统计特征的提示工程,我们创建了受访者的数字孪生体,并分析不同LLMs在复现真实行为上的表现。结果显示,部分模型无法反映真实决策模式,例如错误预测全民疫苗接受度。尽管Llama 3更准确捕捉了种族与收入间的差异,但仍引入了原始UAS数据中不存在的偏见。该研究揭示了生成式代理在行为研究中的潜力,同时强调了来自模型和提示策略的双重偏见风险。
原文摘要 · Abstract (English)
Generative agents have been increasingly used to simulate human behaviour in silico, driven by large language models (LLMs). These simulacra serve as sandboxes for studying human behaviour without compromising privacy or safety. However, it remains unclear whether such agents can truly represent real individuals. This work compares survey data from the Understanding America Study (UAS) on healthcare decision-making with simulated responses from generative agents. Using demographic-based prompt engineering, we create digital twins of survey respondents and analyse how well different LLMs reproduce real-world behaviours. Our findings show that some LLMs fail to reflect realistic decision-making, such as predicting universal vaccine acceptance. However, Llama 3 captures variations across race and Income more accurately but also introduces biases not present in the UAS data. This study highlights the potential of generative agents for behavioural research while underscoring the risks of bias from both LLMs and prompting strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。