arXiv:2512.05647cs.CL2025-12

构建100万份希腊政府决策数据集,助力公共政策分析与智能问答。

A Greek Government Decisions Dataset for Public-Sector Analysis and Insight

  • 从官方平台提取百万级政府决策文本,支持机器读取与复现。
  • 设计RAG任务验证大模型在公文中的检索与推理能力,准确率超基准。
  • 适用于法律/政务领域模型预训练,适合政策研究与AI可解释性探索。

本文发布一个来自希腊国家透明度平台Diavgeia的开放、可机读的政府决策语料库,包含100万份经高质量文本提取的决策文件,以Markdown格式提供原始内容,并附完整可复现的抽取流程。除核心数据集外,我们开展定性分析,识别通用表述模式,并构建了基于代表性问题的检索增强生成(RAG)任务:设计高质量问题与答案,评估基线RAG系统在公共决策文档中的检索与推理表现。实验表明,大规模公共部门语料可有效支撑结构化信息获取与透明度提升,且能模拟交互式政策问答助手。由于其规模、质量与领域覆盖广度,该语料库可作为法律与政府专用语言模型(LM/LLM)的预训练或微调资源,亦可推动领域适应、知识引导生成与可解释人工智能等新方法的发展。最后,我们讨论局限性并开放数据与代码。

原文摘要 · Abstract (English)

We introduce an open, machine-readable corpus of Greek government decisions sourced from the national transparency platform Diavgeia. The resource comprises 1 million decisions, featuring and high-quality raw text extracted from PDFs. It is released with raw extracted text in Markdown format, alongside a fully reproducible extraction pipeline. Beyond the core dataset, we conduct qualitative analyses to explore boilerplate patterns and design a retrieval-augmented generation (RAG) task by formulating a set of representative questions, creating high-quality answers, and evaluating a baseline RAG system on its ability to retrieve and reason over public decisions. This evaluation demonstrates the potential of large-scale public-sector corpora to support advanced information access and transparency through structured retrieval and reasoning over governmental documents, and highlights how such a RAG pipeline could simulate a chat-based assistant capable of interactively answering questions about public decisions. Due to its scale, quality, and domain coverage, the corpus can also serve as high-value pre-training or fine-tuning material for new Language Models (LMs) and Large Language Models (LLMs) respectively, including specialized models for legal and governmental domains, and as a foundation for novel approaches in domain adaptation, knowledge-grounded generation, and explainable AI. Finally, we discuss limitations, outline future directions, and make both the data and the code accessible.

政府数据RAG语言模型公共政策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。