13个大模型在新闻摘要中偏见表现不一,中等规模模型反而更公平。
When Bigger Isn't Better: A Comprehensive Fairness Evaluation of Political Bias in Multi-News Summarisation

- 用政治倾向标注数据集评估13个模型的公平性,发现大模型未必更公平。
- 中等规模模型在5种公平性指标上均优于更大模型,效率更高。
- 实体情感偏见最难纠正,提示工程效果因模型而异,需针对性干预。
多文档新闻摘要系统日益用于处理海量每日新闻内容,跨不同政治立场的公平性至关重要。然而,这些系统可能通过观点代表性不均、特定视角过度强调以及少数声音系统性缺失表现出政治偏见。本研究使用包含政治倾向标签的完整新闻文章数据集FairNews,对13个大型语言模型(LLMs)在五种公平性度量下的表现进行综合评估,考察其对不同政治倾向来源的处理能力,并检验多种去偏策略(包括基于提示和基于判官的方法)。结果挑战了‘模型越大越公平’的假设,发现中等规模模型在公平性与效率间取得最佳平衡。提示式去偏策略效果高度依赖模型,而实体情感维度的偏见最为顽固,所有干预均未能有效缓解。研究表明,多文档新闻摘要的公平性需多维度评估框架与面向架构的针对性去偏方法,而非单纯扩大模型规模。
原文摘要 · Abstract (English)
Multi-document news summarisation systems are increasingly adopted for their convenience in processing vast daily news content, making fairness across diverse political perspectives critical. However, these systems can exhibit political bias through unequal representation of viewpoints, disproportionate emphasis on certain perspectives, and systematic underrepresentation of minority voices. This study presents a comprehensive evaluation of such bias in multi-document news summarisation using FairNews, a dataset of complete news articles with political orientation labels, examining how large language models (LLMs) handle sources with varying political leanings across 13 models and five fairness metrics. We investigate both baseline model performance and effectiveness of various debiasing interventions, including prompt-based and judge-based approaches. Our findings challenge the assumption that larger models yield fairer outputs, as mid-sized variants consistently outperform their larger counterparts, offering the best balance of fairness and efficiency. Prompt-based debiasing proves highly model dependent, while entity sentiment emerges as the most stubborn fairness dimension, resisting all intervention strategies tested. These results demonstrate that fairness in multi-document news summarisation requires multi-dimensional evaluation frameworks and targeted, architecture-aware debiasing rather than simply scaling up.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。