用大模型自动生成符合瑞士法规的可持续采购标准,省时且可审计。
Generating and Evaluating Sustainable Procurement Criteria for the Swiss Public Sector using In-Context Prompting with Large Language Models
- 基于上下文提示技术,结合多模型后端生成采购标准。
- 自动生成的标准与官方指南一致,人工撰写工作量大幅减少。
- 适合公共部门采购人员及政策制定者参考使用。
公共采购指政府、市镇及公共资金机构采购商品和服务的过程。瑞士法律要求在招标评估中融入生态、社会和经济可持续性要求,以具体可验证的标准形式呈现。但将高层次可持续性法规转化为具体、可验证且行业特定的采购标准(如选择标准、评标标准和技术规范)仍需大量领域专业知识和手动操作,耗时且易出错。本文提出一个可配置的、基于大语言模型(LLM)辅助的流水线系统,用于支持瑞士不同采购领域的可持续采购标准目录的系统化生成与评估。该系统整合了上下文提示、可互换的LLM后端和自动化输出验证,实现跨行业的可审计标准生成。作为概念验证,我们使用瑞士政府和欧盟委员会发布的官方可持续性指南作为结构化参考文档进行实例化。通过自动质量检查(包括基于LLM的评估组件)和专家对比人工标注的黄金标准进行评估,结果表明该流水线显著降低手动起草工作量,同时生成的标准与官方指南保持一致。我们还讨论了系统局限性、失败模式及部署中的设计权衡,强调将生成式AI融入公共部门软件流程的关键考量。
原文摘要 · Abstract (English)
Public procurement refers to the process by which public sector institutions, such as governments, municipalities, and publicly funded bodies, acquire goods and services. Swiss law requires the integration of ecological, social, and economic sustainability requirements into tender evaluations in the format of criteria that have to be fulfilled by a bidder. However, translating high-level sustainability regulations into concrete, verifiable, and sector-specific procurement criteria (such as selection criteria, award criteria, and technical specifications) remains a labor-intensive and error-prone manual task, requiring substantial domain expertise in several groups of goods and services and considerable manual effort. This paper presents a configurable, LLM-assisted pipeline that is presented as a software supporting the systematic generation and evaluation of sustainability-oriented procurement criteria catalogs for Switzerland. The system integrates in-context prompting, interchangeable LLM backends, and automated output validation to enable auditable criteria generation across different procurement sectors. As a proof of concept, we instantiate the pipeline using official sustainability guidelines published by the Swiss government and the European Commission, which are ingested as structured reference documents. We evaluate the system through a combination of automated quality checks, including an LLM-based evaluation component, and expert comparison against a manually curated gold standard. Our results demonstrate that the proposed pipeline can substantially reduce manual drafting effort while producing criteria catalogs that are consistent with official guidelines. We further discuss system limitations, failure modes, and design trade-offs observed during deployment, highlighting key considerations for integrating generative AI into public sector software workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。