用参与式预算测试LLM的资源分配能力,发现提示词设计影响关键结果。
LLMs for Resource Allocation: A Participatory Budgeting Approach to Inferring Preferences
- 用参与式预算作为场景和动态评估基准,测试LLM决策能力。
- 在预算约束下,贪心与优化策略表现接近最优解,误差小于5%。
- 能从自然语言中推断偏好,适合用于无明确投票的机制设计场景。
大型语言模型(LLMs)日益被期望承担复杂决策任务,但其在结构化资源分配方面的能力仍缺乏探索。现有评估基准因数据污染和静态性难以有效衡量其推理能力。本文提出一种双用途框架,利用参与式预算(Participatory Budgeting, PB)作为(i)LLM资源分配的实际应用场景,以及(ii)动态评估其推理能力的自适应基准。我们采用三种提示策略:贪心选择、直接优化和类爬山优化迭代,让LLM在预算等可行性约束下选择项目子集。将模型分配结果与效用最大化基准(utility-maximizing oracle)对比,结果显示提示设计显著影响性能。同时,我们测试了LLM能否从自然语言选民输入或元数据中推断出结构化偏好,而无需显式投票。通过比较基于推断偏好与真实投票的分配结果,评估其从非结构化输入中提取偏好的能力。结果表明,提示工程对模型表现至关重要,且LLM在处理非结构化输入的机制设计中具有潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly expected to handle complex decision-making tasks, yet their ability to perform structured resource allocation remains underexplored. Evaluating their reasoning is also difficult due to data contamination and the static nature of existing benchmarks. We present a dual-purpose framework leveraging Participatory Budgeting (PB) both as (i) a practical setting for LLM-based resource allocation and (ii) an adaptive benchmark for evaluating their reasoning capabilities. We task LLMs with selecting project subsets under feasibility (e.g., budget) constraints via three prompting strategies: greedy selection, direct optimization, and a hill-climbing-inspired refinement. We benchmark LLMs' allocations against a utility-maximizing oracle. Interestingly, we also test whether LLMs can infer structured preferences from natural-language voter input or metadata, without explicit votes. By comparing allocations based on inferred preferences to those from ground-truth votes, we evaluate LLMs' ability to extract preferences from open-ended input. Our results underscore the role of prompt design and show that LLMs hold promise for mechanism design with unstructured inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。