arXiv:2508.08193cs.CYcs.AI2025-08AAAI被引 6

测试大模型在救助资源分配中的判断能力,发现其结果极不稳定。

Street-Level AI: Are Large Language Models Ready for Real-World Judgments?

  • 用真实求助数据测试大模型优先级判断
  • 模型结果在不同运行间差异巨大,与评分系统不一致
  • 虽与普通人判断有相似性,但不适合直接用于社会决策

近期大量研究关注大模型做出“道德判断”的伦理与社会影响。多数工作聚焦于对齐人类判断或群体公平性,但大模型最可能的现实用途是替代街头级公务员——即决定稀缺社会资源分配或福利审批的基层人员。该领域已有长期关于本地正义原则如何影响优先级机制的历史。本文检验大模型判断与人类判断、以及当前无家可归者资源分配中使用的社会政治确定的脆弱性评分系统之间的对齐程度。关键的是,我们使用真实需求数据(通过本地大模型保持保密)进行分析。结果表明,大模型的优先级判断在多个层面极不一致:同一模型不同运行之间、不同模型之间,以及与脆弱性评分系统之间均存在显著差异。同时,大模型在成对比较测试中表现出与普通人的定性一致性。研究质疑当前大模型在高风险社会决策中直接应用的准备度。

原文摘要 · Abstract (English)

A surge of recent work explores the ethical and societal implications of large-scale AI models that make "moral" judgments. Much of this literature focuses either on alignment with human judgments through various thought experiments or on the group fairness implications of AI judgments. However, the most immediate and likely use of AI is to help or fully replace the so-called street-level bureaucrats, the individuals deciding to allocate scarce social resources or approve benefits. There is a rich history underlying how principles of local justice determine how society decides on prioritization mechanisms in such domains. In this paper, we examine how well LLM judgments align with human judgments, as well as with socially and politically determined vulnerability scoring systems currently used in the domain of homelessness resource allocation. Crucially, we use real data on those needing services (maintaining strict confidentiality by only using local large models) to perform our analyses. We find that LLM prioritizations are extremely inconsistent in several ways: internally on different runs, between different LLMs, and between LLMs and the vulnerability scoring systems. At the same time, LLMs demonstrate qualitative consistency with lay human judgments in pairwise testing. Findings call into question the readiness of current generation AI systems for naive integration in high-stakes societal decision-making.

大模型评估社会公平决策系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。