arXiv:2505.02252cs.CL2025-05AAAI被引 1

用去偏微调让大模型在不同地区背景下更公平地识别仇恨言论。

Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models

  • 通过惩罚有无地域上下文时的分类不一致来去偏。
  • 微调后模型在个性化与无上下文场景下表现均提升。
  • 揭示了个性化信息可能放大地域偏见,值得安全敏感应用关注。

商用大型语言模型(LLMs)最近引入了记忆功能以提供个性化响应,这些记忆会保留用户人口统计和个体特征等信息,使模型能根据个人数据调整行为。然而,将个性化信息融入上下文的影响尚未充分评估,尤其是在敏感话题上。本文研究了多种先进LLM在不同个性化场景下的表现,聚焦于仇恨言论检测。我们通过让模型假设特定国家身份并使用不同语言进行检测,发现上下文个性化显著影响其对仇恨言论的判断。为缓解此类偏差,我们通过对模型进行微调,惩罚在有/无国家或语言上下文时出现的不一致分类。经优化的模型在个性化情境及无上下文时均表现出更好性能。

原文摘要 · Abstract (English)

Commercial Large Language Models (LLMs) have recently incorporated memory features to deliver personalised responses. This memory retains details such as user demographics and individual characteristics, allowing LLMs to adjust their behaviour based on personal information. However, the impact of integrating personalised information into the context has not been thoroughly assessed, leading to questions about its influence on LLM behaviour. Personalisation can be challenging, particularly with sensitive topics. In this paper, we examine various state-of-the-art LLMs to understand their behaviour in different personalisation scenarios, specifically focusing on hate speech. We prompt the models to assume country-specific personas and use different languages for hate speech detection. Our findings reveal that context personalisation significantly influences LLMs' responses in this sensitive area. To mitigate these unwanted biases, we fine-tune the LLMs by penalising inconsistent hate speech classifications made with and without country or language-specific context. The refined models demonstrate improved performance in both personalised contexts and when no context is provided.

大模型去偏仇恨言论个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。