用大模型生成政策改进建议,首次建立评估基准
Report-based Recommendations for Policy Making and Agency Operations: Dataset and LLM Evaluation
- 基于报告内容生成组织政策改进建议
- 顶尖大模型能有效提炼关键问题与经验教训
- 适合政策研究、政府机构及企业战略部门参考
大型语言模型(LLMs)在文本生成任务中应用广泛,其生成能力使它们有望为政策制定或机构运营提供有价值见解。本文提出一项新任务:基于报告内容生成可指导未来行动和改进的建议,特别聚焦于私营与公共组织的政策优化。我们首次构建了该任务的基准数据集与系统性评估框架,区别于传统的产品或用户推荐系统,本任务旨在从报告结论中提炼出政策改进建议。实验结果表明,当前最先进的大模型具备识别并反思关键问题与学习要点的能力,展现出在政策支持领域的潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are extensively used in text generation tasks. These generative capabilities bring us to a point where LLMs could potentially provide useful insights in policy making or agency operations. In this paper, we introduce a new task consisting of generating recommendations which can be used to inform future actions and improvements of agencies work within private and public organisations. In particular, we present the first benchmark and coherent evaluation for developing recommendation systems to inform organisation policies. This task is clearly different from usual product or user recommendation systems, but rather aims at providing a basis to suggest policy improvements based on the conclusions drawn from reports. Our results demonstrate that state-of-the-art LLMs have the potential to emphasize and reflect on key issues and learning points within generated recommendations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。