用极化敏感方法评估大模型在敏感话题上的偏见,发现模型倾向性差异显著。
BIPOLAR: Polarization-based granular framework for LLM bias evaluation
- 基于情感极化指标与合成平衡数据集,构建可复用的偏见评估框架。
- 在俄乌战争语境下,各模型对乌克兰普遍持更积极态度,但类别间差异明显。
- 适用于多种敏感议题,支持自动化生成与细粒度分析,适合安全与伦理研究者。
大语言模型在处理政治言论、性别认同、民族关系或国家刻板印象等敏感话题时,常表现出偏见。尽管偏见检测与缓解技术已有进展,但仍存在未充分探索的挑战。本文提出一种可复用、细粒度且主题无关的框架,用于评估大语言模型(开源与闭源)在极化相关偏见方面的表现。该方法结合极化敏感的情感度量与人工合成的冲突类陈述平衡数据集,使用预定义的语义类别。以俄乌战争为例,构建了合成数据集,并评估了 Llama-3、Mistral、GPT-4、Claude 3.5 与 Gemini 1.0 等模型。除总体偏见分数外,框架揭示了不同语义类别间显著差异,暴露了模型行为模式的多样性。提示词修改后的适应性分析进一步显示,模型倾向于预设语言和国籍表达。整体上,该框架支持自动化数据生成与细粒度偏见评估,适用于多种极化驱动场景,且与其他评估策略正交。
原文摘要 · Abstract (English)
Large language models (LLMs) are known to exhibit biases in downstream tasks, especially when dealing with sensitive topics such as political discourse, gender identity, ethnic relations, or national stereotypes. Although significant progress has been made in bias detection and mitigation techniques, certain challenges remain underexplored. This study proposes a reusable, granular, and topic-agnostic framework to evaluate polarisation-related biases in LLM (both open-source and closed-source). Our approach combines polarisation-sensitive sentiment metrics with a synthetically generated balanced dataset of conflict-related statements, using a predefined set of semantic categories. As a case study, we created a synthetic dataset that focusses on the Russia-Ukraine war, and we evaluated the bias in several LLMs: Llama-3, Mistral, GPT-4, Claude 3.5, and Gemini 1.0. Beyond aggregate bias scores, with a general trend for more positive sentiment toward Ukraine, the framework allowed fine-grained analysis with considerable variation between semantic categories, uncovering divergent behavioural patterns among models. Adaptation to prompt modifications showed further bias towards preconceived language and citizenship modification. Overall, the framework supports automated dataset generation and fine-grained bias assessment, is applicable to a variety of polarisation-driven scenarios and topics, and is orthogonal to many other bias-evaluation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。