arXiv:2601.14944cs.CL2026-01ACL

构建民主议事数据集,用小模型实现自动清理与结构化分析。

The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations

  • 提出语料澄清框架,将杂乱议事文本转为可分析的论证单元。
  • 人工标注1231条法国全民大辩论数据,含2285个论证单元。
  • 小模型表现媲美大模型,适合本地部署的透明民主分析。

大型语言模型在现代自然语言处理中广泛应用,尽管其可用于在线审议或大规模公民咨询等民主活动文本,但其作为分析工具的使用仍引发伦理争议。本文旨在实现两个目标:(a) 开发资源以在实践层面标准化公共论坛中的公民贡献,使其更适用于主题建模与政治分析;(b) 研究小型开源权重语言模型(可在本地运行、资源消耗低)能否可靠完成此类标准化。为此,我们提出语料澄清(Corpus Clarification)框架,将大规模咨询数据中的噪声文本、多主题内容转化为结构化、自包含的论证单元,便于下游分析。我们构建了GDN-CC数据集,包含1,231条法国全民大辩论(Grand Débat National)的贡献,共2,285个经过论证结构标注并手动澄清的单元。实验表明,微调后的小型语言模型在复现标注方面表现不逊于甚至优于大型模型,并在意见聚类任务中验证了其可用性。最后,我们发布了规模达24万条的自动化标注语料GDN-CC-large,为迄今最大规模的民主咨询标注数据集。

原文摘要 · Abstract (English)

LLMs are ubiquitous in modern NLP, and while their applicability extends to texts produced for democratic activities such as online deliberations or large-scale citizen consultations, ethical questions have been raised for their usage as analysis tools. We continue this line of research with two main goals: (a) to develop resources that can help standardize citizen contributions in public forums at the pragmatic level, and make them easier to use in topic modeling and political analysis; (b) to study how well this standardization can reliably be performed by small, open-weights LLMs, i.e. models that can be run locally and transparently with limited resources. Accordingly, we introduce Corpus Clarification as a preprocessing framework for large-scale consultation data that transforms noisy, multi-topic contributions into structured, self-contained argumentative units ready for downstream analysis. We present GDN-CC, a manually-curated dataset of 1,231 contributions to the French Grand Débat National, comprising 2,285 argumentative units annotated for argumentative structure and manually clarified. We then show that finetuned Small Language Models match or outperform LLMs on reproducing these annotations, and measure their usability for an opinion clustering task. We finally release GDN-CC-large, an automatically annotated corpus of 240k contributions, the largest annotated democratic consultation dataset to date.

民主咨询数据标注小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。