arXiv:2410.15471cs.AIcs.LG2024-10

测试大模型在再犯预测中的表现,发现其不如人类和传统模型可靠。

Generative Models, Humans, Predictive Models: Who Is Worse at High-Stakes Decision Making?

  • 对比大模型、人类与预测模型在再犯预测任务中的决策一致性。
  • 引入干扰信息后,大模型决策稳定性下降,且部分优化技术反而加剧偏差。
  • 提醒当前大模型不适合高风险决策场景,尤其在有偏数据下更不可靠。

尽管已有明确警告,大型生成式语言模型(LMs)仍被用于原本由预测模型或人类完成的高风险决策任务。本文在再犯预测这一高风险任务中测试了三种闭源和开源的大模型,不仅评估其准确性,还分析其与人类判断及现有预测模型的一致性(后者本身存在不完美、噪声和偏见)。实验考察了提供不同类型信息(如照片等干扰信息)对模型决策的影响,并对旨在提升准确率或缓解偏见的技术进行压力测试,发现部分方法产生了意料之外的负面影响。结果提供了定量证据,支持当前大模型并不适合作为高风险决策工具的判断。

原文摘要 · Abstract (English)

Despite strong advisory against it, large generative models (LMs) are already being used for decision making tasks that were previously done by predictive models or humans. We put popular LMs to the test in a high-stakes decision making task: recidivism prediction. Studying three closed-access and open-source LMs, we analyze the LMs not exclusively in terms of accuracy, but also in terms of agreement with (imperfect, noisy, and sometimes biased) human predictions or existing predictive models. We conduct experiments that assess how providing different types of information, including distractor information such as photos, can influence LM decisions. We also stress test techniques designed to either increase accuracy or mitigate bias in LMs, and find that some to have unintended consequences on LM decisions. Our results provide additional quantitative evidence to the wisdom that current LMs are not the right tools for these types of tasks.

大模型评估再犯预测决策可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。