构建议会演讲生成评测框架,提升大模型的政治立场一致性
ParliaBench: An Evaluation and Benchmarking Framework for LLM-Generated Parliamentary Speech
- 用英国议会数据集训练并设计多维评估体系
- 微调后模型在语言质量与政治真实性上显著提升
- 提出两个新指标量化政党立场与意识形态分布
议会演讲生成对大语言模型提出了超越常规文本生成的挑战。除语言质量外,还需具备政治真实性和意识形态一致性。现有模型缺乏议会场景专门训练,评估方法也仅依赖通用NLP指标。为此,我们提出ParliaBench,一个针对议会演讲生成的基准评测框架。构建了来自英国议会的演讲数据集以支持系统性模型训练。设计包含计算指标与大模型评判的评估体系,从语言质量、语义连贯性、政治真实性三个维度衡量生成效果。提出两项基于嵌入的新指标:政治光谱对齐(Political Spectrum Alignment)和政党对齐(Party Alignment),用于量化意识形态定位。对五种大语言模型进行微调,生成28,000篇演讲,并使用该框架进行评估,对比基线与微调模型表现。结果表明,微调在多数指标上带来统计显著提升,新指标在政治维度上表现出强区分能力。
原文摘要 · Abstract (English)
Parliamentary speech generation presents specific challenges for large language models beyond standard text generation tasks. Unlike general text generation, parliamentary speeches require not only linguistic quality but also political authenticity and ideological consistency. Current language models lack specialized training for parliamentary contexts, and existing evaluation methods focus on standard NLP metrics rather than political authenticity. To address this, we present ParliaBench, a benchmark for parliamentary speech generation. We constructed a dataset of speeches from UK Parliament to enable systematic model training. We introduce an evaluation framework combining computational metrics with LLM-as-a-judge assessments for measuring generation quality across three dimensions: linguistic quality, semantic coherence, and political authenticity. We propose two novel embedding-based metrics, Political Spectrum Alignment and Party Alignment, to quantify ideological positioning. We fine-tuned five large language models (LLMs), generated 28k speeches, and evaluated them using our framework, comparing baseline and fine-tuned models. Results show that fine-tuning produces statistically significant improvements across the majority of metrics and our novel metrics demonstrate strong discriminative power for political dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。