检测并缓解大模型在议会辩论摘要中的代表性偏差。
Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
- 设计可归因的评估框架,量化发言者被包含或误述的程度。
- 发现中段发言、非英语及左翼政党更易被忽略或歪曲。
- 提出分层摘要法有效缓解位置偏差,适合政策与伦理研究者。
大型语言模型(LLMs)用于总结议会会议内容,提升了民主参与的可及性,但其作为政治信息中介时存在公平性问题。本文评估了5个主流LLM(含专有与开源模型)对欧洲议会全体会议辩论的摘要表现,提出一种归因感知的评估框架,衡量发言者的代表性与误述情况。结果表明:发言顺序(中间发言被系统性排除)、语言(非英语发言者代表性较低)及政治立场(左翼政党受益更多)均导致显著偏差。我们进一步区分了遗漏偏差(inclusion bias)与虚构偏差(hallucination bias),发现提示工程无法缓解偏差。提出分层摘要方法,将任务分解为提取与聚合两步,显著改善所有模型的位置偏差。研究强调需采用领域敏感的评估指标与伦理监督,以保障多语言民主应用的公正性。
原文摘要 · Abstract (English)
The The use of Large language models (LLMs) to summarise parliamentary proceedings presents a promising means of increasing the accessibility of democratic participation. However, as these systems increasingly mediate access to political information -- filtering and framing content before it reaches users -- there are important fairness considerations to address. In this work, we evaluate 5 LLMs (both proprietary and open-weight) in the summarisation of plenary debates from the European Parliament to investigate the representational biases that emerge in this context. We develop an attribution-aware evaluation framework to measure speaker-level inclusion and mis-representation in debate summaries. Across all models and experiments, we find that speakers are less accurately represented in the final summary on the basis of (i) their speaking-order (speeches in the middle of the debate were systematically excluded), (ii) language spoken (non-English speakers were less faithfully represented), and (iii) political affiliations (better outcomes for left-of-centre parties). We further show how biases in these contexts can be decomposed to distinguish inclusion bias (systematic omission) from hallucination bias (systematic misrepresentation), and explore the effect of different mitigation strategies. Prompting strategies do not affect these biases. Instead, we propose a hierarchical summarisation method that decomposes the task into simpler extraction and aggregation steps, which we show significantly improves the positional/speaking-order bias across all models. These findings underscore the need for domain-sensitive evaluation metrics and ethical oversight in the deployment of LLMs for multilingual democratic applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。