首次量化大模型新闻摘要的框架偏差,发现其偏见程度高于人类撰写的参考文本。
Frame In, Frame Out: Measuring Framing Bias in LLM-Generated News Summaries
- 构建首个大规模基准FIFO,结合人工标注数据评估生成摘要的框架偏差。
- 27个模型测试显示,大模型摘要的框架偏差率普遍高于人类参考文本,科学类更高。
- 适用于关注信息偏见、内容安全与可信赖性的研究者和平台方。
新闻标题与摘要通过选择性强调与删减来塑造事件解读,这一现象称为框架效应。大型语言模型现常用于生成此类内容,但现有评估体系大多忽略此维度。我们提出帧内-帧外(Frame In, Frame Out, FIFO),首个基于XSum数据集的大规模基准,用于测量大模型生成摘要中的框架存在性。FIFO包含15,499个评审标注样本与320个专家标注实例(κ=0.61),用于验证与校准模型标注。利用FIFO,我们分析了27个摘要模型的框架率。结果表明,大模型生成摘要的校准框架率通常高于人类撰写的参考文本,且在不同主题与训练方式间存在显著差异,尤其在科学与公共卫生类摘要中更为突出。研究确立了框架偏差作为摘要质量中被忽视但关键的维度。
原文摘要 · Abstract (English)
News headlines and summaries shape how events are interpreted through selective emphasis and omission, a phenomenon commonly referred to as framing. Large language models are now routinely used to generate such content, yet existing evaluation frameworks largely overlook this dimension. We introduce Frame In, Frame Out (FIFO), the first large-scale benchmark for measuring framing presence in LLM-generated news summaries, grounded in the widely used XSum dataset. FIFO combines 15,499 jury-annotated examples with 320 expert-labeled instances ($κ= 0.61$) to validate and calibrate model-based annotations. Using FIFO, we analyze measured framing rates across 27 summarization models. We find that LLM-generated summaries often exhibit higher calibrated framing rates than human-written references, with substantial variation across topics and training regimes, including elevated rates in scientific and public health summaries. Our results establish framing as an underexplored and consequential dimension of summarization quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。