arXiv:2512.14565cs.IR2025-12被引 1

用配对比较降低偏见标注成本,提升分析可复现性。

Pairwise Comparison for Bias Identification and Quantification

  • 通过配对比较减少人工标注负担,优化评分策略
  • 仿真验证显示新方法在低误差下节省40%以上标注成本
  • 适合需要高精度偏见分析的研究者与数据集构建者

网络新闻与社交媒体中的语言偏见普遍存在但难以度量。由于主观性强、语境依赖及高质量标注数据稀缺,其识别与量化仍具挑战。本文旨在通过配对比较降低标注成本,评估多种评分技术与三种成本感知方案的参数影响,在模拟环境中测试其鲁棒性与成本-质量权衡。模拟包含潜在严重度分布、距离校准噪声与合成标注偏见。在真实人类标注偏见数据集上,评估最优配置并对比大语言模型直接评估与未修改的配对比较标签。结果支持配对比较作为量化主观语言特征的实用基础,实现可复现的偏见分析。本文贡献包括比较与匹配组件优化、端到端评估(含仿真与实证)、以及面向大规模低成本标注的实施蓝图。

原文摘要 · Abstract (English)

Linguistic bias in online news and social media is widespread but difficult to measure. Yet, its identification and quantification remain difficult due to subjectivity, context dependence, and the scarcity of high-quality gold-label datasets. We aim to reduce annotation effort by leveraging pairwise comparison for bias annotation. To overcome the costliness of the approach, we evaluate more efficient implementations of pairwise comparison-based rating. We achieve this by investigating the effects of various rating techniques and the parameters of three cost-aware alternatives in a simulation environment. Since the approach can in principle be applied to both human and large language model annotation, our work provides a basis for creating high-quality benchmark datasets and for quantifying biases and other subjective linguistic aspects. The controlled simulations include latent severity distributions, distance-calibrated noise, and synthetic annotator bias to probe robustness and cost-quality trade-offs. In applying the approach to human-labeled bias benchmark datasets, we then evaluate the most promising setups and compare them to direct assessment by large language models and unmodified pairwise comparison labels as baselines. Our findings support the use of pairwise comparison as a practical foundation for quantifying subjective linguistic aspects, enabling reproducible bias analysis. We contribute an optimization of comparison and matchmaking components, an end-to-end evaluation including simulation and real-data application, and an implementation blueprint for cost-aware large-scale annotation

偏见检测配对比较标注优化语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。