arXiv:2605.28848cs.CLcs.AI2026-05

实时检测大模型对不同群体的新闻表述差异,发现政策类提示影响最大。

GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models

论文配图:GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models
图 1 · 摘自论文原文
  • 构建动态新闻流评估协议,覆盖42类身份标签与7种提示类型。
  • 政策/行动类提示引发最强语义变化,情感差异相对稳定。
  • 适合关注模型偏见审计与实时安全监测的研究者使用。

部署中的语言模型面临非平稳环境:模型版本、检索层、安全系统及真实输入随时间变化。静态偏见基准仍有价值,但无法揭示模型如何为不同目标受众描述新出现事件。本文提出GPF-LIVENEWS,一种用于审计开放性大模型输出中群体相关表述的流式评估协议与基准快照。该协议扩展了来自BBC/Reuters的最新新闻主播内容,涵盖42个身份标签和7种提示类别,并通过语义敏感性和情感差异信号评估响应包。在12次监控运行与23个托管模型的试点中,政策/行动类提示产生最强语义变动,而情感差异在各维度与提示族间较为平缓。发布的成果包含文章元数据、提示模板、实例化提示、模型输出元数据、评分表、文档及复现脚本。所有得分均视为观察窗口内的审计信号,供人工审查,而非永久公平性排名或有害偏见的直接证据。

原文摘要 · Abstract (English)

Deployed language models are evaluated in a non-stationary environment: model versions, retrieval layers, safety systems, and real-world inputs all change over time. Static bias benchmarks remain useful, but they do not show how models frame newly emerging events for different prompted audiences. We introduce GPF-LIVENEWS, a streaming evaluation protocol and benchmark snapshot for auditing group-conditioned framing in open-ended LLM outputs. The protocol expands fresh BBC/Reuters news anchors across 42 identity labels and seven prompt families, then evaluates response bundles using semantic-sensitivity and sentiment-disparity signals. In a pilot over 12 monitoring runs and 23 hosted models, Policy/Action prompts produce the strongest semantic movement, while sentiment variation is flatter across dimensions and prompt families. The released artifact includes article metadata, prompt templates, instantiated prompts, model-output metadata, score tables, documentation, and reproduction scripts. We interpret all scores as observed-window audit signals for human review, not as permanent fairness rankings or direct proof of harmful bias.

模型审计新闻生成偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。