用大模型评估救助无家可归者的政策,发现其建议与专家有差异但具参考价值。
What Would an LLM Do? Evaluating Large Language Models for Policymaking to Alleviate Homelessness
- 构建四城市政策场景基准,结合人类发展能力框架
- 大模型建议与专家意见存在差异,但能生成有潜力的政策方案
- 适合政策研究者、城市规划者及人本导向决策者参考
大型语言模型(LLMs)在高风险领域中的应用日益广泛。它们能够捕捉不断变化的社会背景并生成合理的情景,因此在社会政策制定中具有潜力。本文评估了大模型在缓解无家可归问题上的政策建议是否与领域专家一致(以及彼此之间的一致性),该问题影响全球超过1.5亿人。我们构建了一个新颖的基准测试,涵盖四个城市的决策场景,政策选项基于人类发展能力方法论框架。同时提出自动化流程,将政策建议连接至一个地点的基于代理的模型,并比较大模型与专家推荐政策的社会影响。初步分析显示,不同大模型的政策建议与本地专家存在差异,但若结合负责任的防护机制、情境校准和本地专业知识,大模型仍可能为政策制定提供有益见解。本研究将能力方法论转化为计算框架,为以人类尊严为核心的无家可归者救助政策提供了新洞见。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly being adopted in high-stakes domains. Their potential to encode evolving social contexts and to generate plausible scenarios position them as promising tools in social policymaking. This article evaluates whether LLMs are aligned with domain experts (and among themselves) on policy recommendations to alleviate homelessness - a challenge affecting over 150 million people worldwide. We develop a novel benchmark comprised of decision scenarios across four cities, with policy choices that are grounded in the conceptual framework of the Capability Approach for human development. We also present an automated pipeline that connects the policies to an agent-based model in one location, and compare the social impact of the policies recommended by LLMs to those recommended by experts. Our exploratory analysis reveals variation across LLMs in their policy recommendations compared to local experts, yet suggests potential benefits of the use of LLMs to provide insights for policymaking, if paired with responsible guardrails, contextual calibrations, and local domain expertise. Our work operationalizes the Capability Approach in a computational framework and provides new insights on homelessness alleviation policymaking with a focus on human dignity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。