构建首个大规模人工标注的辩论数据集,评估大模型在政策讨论中的代表性与公平性。
Can AI Truly Represent Your Voice in Deliberations? A Comprehensive Study of Large-Scale Opinion Aggregation with LLMs
- 创建包含3000人参与的十类议题数据集,支持多维度人工评分。
- 提出专用评判模型DeliberationJudge,比大模型更贴近人类判断。
- 发现主流大模型普遍存在少数意见被忽视问题,适合政策AI评估者使用。
大规模公众讨论产生数千条自由文本观点,需提炼为具代表性和中立性的摘要以供政策参考。尽管大语言模型(LLMs)在生成此类摘要方面表现出潜力,但其可能低估少数派观点且受输入顺序影响,引发高风险场景下的公平性问题。现有研究常依赖大模型自身作为评价标准,但其与人类判断对齐度较差。为此,本文构建了DeliberationBank——一个大规模人工基准数据集:(1) 覆盖10个议题、由3000名参与者贡献的原始意见;(2) 由4500名参与者在代表性、信息量、中立性、政策支持度四个维度标注的摘要评价数据。基于此,我们训练了专用评判模型DeliberationJudge(基于DeBERTa微调),可从个体视角评估摘要质量。该模型在效率与人类判断一致性上均优于多种大模型评判器。利用该框架,我们评估了18个大模型,揭示其在汇总过程中普遍存在的少数派观点遗漏问题。本框架为政策决策中的辩论摘要评估提供了可扩展、可靠的评测路径,有助于提升AI系统的代表性与公平性。
原文摘要 · Abstract (English)
Large-scale public deliberations generate thousands of free-form contributions that must be synthesized into representative and neutral summaries for policy use. While LLMs have been shown as a promising tool to generate summaries for large-scale deliberations, they also risk underrepresenting minority perspectives and exhibiting bias with respect to the input order, raising fairness concerns in high-stakes contexts. Studying and fixing these issues requires a comprehensive evaluation at a large scale, yet current practice often relies on LLMs as judges, which show weak alignment with human judgments. To address this, we present DeliberationBank, a large-scale human-grounded dataset with (1) opinion data spanning ten deliberation questions created by 3,000 participants and (2) summary judgment data annotated by 4,500 participants across four dimensions (representativeness, informativeness, neutrality, policy approval). Using these datasets, we train DeliberationJudge, a fine-tuned DeBERTa model that can rate deliberation summaries from individual perspectives. DeliberationJudge is more efficient and more aligned with human judgements compared to a wide range of LLM judges. With DeliberationJudge, we evaluate 18 LLMs and reveal persistent weaknesses in deliberation summarization, especially underrepresentation of minority positions. Our framework provides a scalable and reliable way to evaluate deliberation summarization, helping ensure AI systems are more representative and equitable for policymaking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。