arXiv:2511.14662cs.CL2025-11被引 5

揭示多语言大模型标注偏见根源,提出系统性应对方案

Bias in, Bias out: Annotation Bias in Multilingual Large Language Models

  • 区分指令、标注者和文化三类标注偏见
  • 提出多语言模型差异与文化推断等新检测方法
  • 适合关注公平性与跨文化AI的开发者与研究者

自然语言处理数据集中的标注偏见仍是构建多语言大模型的主要挑战,尤其在文化多样性背景下。任务表述偏差、标注者主观性及文化错配可能扭曲模型输出并加剧社会危害。本文提出一套全面的标注偏见理解框架,区分指令偏见、标注者偏见以及上下文与文化偏见。综述了检测方法(包括标注者间一致性、模型分歧与元数据分析),并强调多语言模型差异与文化推理等新兴技术。进一步提出主动与被动缓解策略,如多样化标注者招募、迭代指南优化及事后模型调整。贡献包括:(1) 标注偏见分类体系;(2) 检测指标整合;(3) 适配多语言场景的集成式偏见缓解方法;(4) 标注流程的伦理分析。这些见解旨在推动更公平、文化契合的多语言大模型标注流程。

原文摘要 · Abstract (English)

Annotation bias in NLP datasets remains a major challenge for developing multilingual Large Language Models (LLMs), particularly in culturally diverse settings. Bias from task framing, annotator subjectivity, and cultural mismatches can distort model outputs and exacerbate social harms. We propose a comprehensive framework for understanding annotation bias, distinguishing among instruction bias, annotator bias, and contextual and cultural bias. We review detection methods (including inter-annotator agreement, model disagreement, and metadata analysis) and highlight emerging techniques such as multilingual model divergence and cultural inference. We further outline proactive and reactive mitigation strategies, including diverse annotator recruitment, iterative guideline refinement, and post-hoc model adjustments. Our contributions include: (1) a typology of annotation bias; (2) a synthesis of detection metrics; (3) an ensemble-based bias mitigation approach adapted for multilingual settings, and (4) an ethical analysis of annotation processes. Together, these insights aim to inform more equitable and culturally grounded annotation pipelines for LLMs.

标注偏见多语言伦理公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。