arXiv:2410.14744cs.CLcs.AI2024-10被引 4

让大模型表达不确定,能减少预测有害行为时的偏见。

Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors

  • 要求模型显式表达不确定性,以降低预测偏差。
  • 在两个社交平台数据集上,不确定性表示显著提升公平性。
  • 无需大量标注数据,即可有效缓解模型偏见,适合安全审核场景。

对话预测任务要求模型预判未完成对话的最终结果,例如在社交媒体内容审核中提前识别潜在有害用户行为,从而实现预防性干预。尽管大型语言模型(LLMs)已被证明在对话预测中表现有效,但其对特定预测目标(如有害行为)可能存在隐含偏见。本文探究模型不确定性表达在缓解此类偏见中的作用,提出三个核心研究问题:1)当要求模型表达不确定性时,其预测准确率如何变化;2)模型偏见是否随不确定性表达而改变;3)能否利用不确定性表示在少量训练数据下消除或减轻偏见。研究针对5个开源语言模型,在2个专为社交媒体内容审核设计的对话预测数据集上进行验证。

原文摘要 · Abstract (English)

Conversation forecasting tasks a model with predicting the outcome of an unfolding conversation. For instance, it can be applied in social media moderation to predict harmful user behaviors before they occur, allowing for preventative interventions. While large language models (LLMs) have recently been proposed as an effective tool for conversation forecasting, it's unclear what biases they may have, especially against forecasting the (potentially harmful) outcomes we request them to predict during moderation. This paper explores to what extent model uncertainty can be used as a tool to mitigate potential biases. Specifically, we ask three primary research questions: 1) how does LLM forecasting accuracy change when we ask models to represent their uncertainty; 2) how does LLM bias change when we ask models to represent their uncertainty; 3) how can we use uncertainty representations to reduce or completely mitigate biases without many training data points. We address these questions for 5 open-source language models tested on 2 datasets designed to evaluate conversation forecasting for social media moderation.

对话预测模型偏见不确定性建模内容审核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。