arXiv:2505.16081cs.CL2025-05被引 1

构建政治偏见数据集,用双轴标注+理由标注实现可解释检测

BiasLab: Toward Explainable Political Bias Detection with Dual-Axis Annotations and Rationale Indicators

  • 双轴标注民主/共和党倾向,每篇配理由说明判断依据
  • 发现模型常误判微妙右倾内容,与人类标注存在对称性偏差
  • 适合研究可解释性NLP、政治偏见检测及人机协同评估

我们提出BiasLab,一个包含300篇政治新闻文章的标注数据集,覆盖多样政治事件与媒体立场。每篇文章由众包工作者在民主党和共和党倾向两个独立维度上打分,并附加理由标注。数据集来自900篇精选文档池,通过针对性工人筛选和试点分析优化标注流程。我们量化了标注者间一致性,分析了与媒体源立场的错位情况,并将标签划分为可解释子集。此外,使用受约束的GPT-4o模拟标注,发现其与人类标注存在镜像偏差,尤其在识别轻微右倾内容时表现不佳。定义了感知漂移预测和理由类型分类两项任务,报告基线性能以展示可解释偏见检测的挑战。丰富的理由标注为政治偏见建模提供可操作的解释,支持开发透明、社会敏感的NLP系统。数据集、标注框架和建模代码已公开,促进人机协同可解释性研究与真实场景下解释有效性评估。

原文摘要 · Abstract (English)

We present BiasLab, a dataset of 300 political news articles annotated for perceived ideological bias. These articles were selected from a curated 900-document pool covering diverse political events and source biases. Each article is labeled by crowdworkers along two independent scales, assessing sentiment toward the Democratic and Republican parties, and enriched with rationale indicators. The annotation pipeline incorporates targeted worker qualification and was refined through pilot-phase analysis. We quantify inter-annotator agreement, analyze misalignment with source-level outlet bias, and organize the resulting labels into interpretable subsets. Additionally, we simulate annotation using schema-constrained GPT-4o, enabling direct comparison to human labels and revealing mirrored asymmetries, especially in misclassifying subtly right-leaning content. We define two modeling tasks: perception drift prediction and rationale type classification, and report baseline performance to illustrate the challenge of explainable bias detection. BiasLab's rich rationale annotations provide actionable interpretations that facilitate explainable modeling of political bias, supporting the development of transparent, socially aware NLP systems. We release the dataset, annotation schema, and modeling code to encourage research on human-in-the-loop interpretability and the evaluation of explanation effectiveness in real-world settings.

政治偏见可解释性双轴标注理由标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。