arXiv:2410.19317cs.CL2024-10ICLR被引 48

构建多轮对话公平性评测基准,揭示大模型在真实对话中偏见累积问题。

FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs

  • 设计三阶段任务体系,覆盖理解、交互与指令权衡中的公平性挑战。
  • 构建10,000条多轮对话数据集,发现当前模型在长对话中更易产生偏见。
  • 推出1,000条高难度测试集,适合评估和改进大模型的真实对话公平性。

基于大语言模型的聊天机器人广泛应用引发公平性担忧。现有公平性评测多聚焦单轮对话,而多轮对话因对话复杂性和偏见累积更具挑战性。本文提出一个全面的多轮对话公平性评测基准 FairMT-Bench,涵盖上下文理解、用户交互与指令权衡三个阶段,每阶段设两个任务。通过整合现有公平性数据集并使用模板构建多轮对话数据集 FairMT-10K(10,000 条)。采用 GPT-4、Llama-Guard-3 偏见分类器与人工验证进行评估。实验表明,在多轮场景下,当前大模型更易生成偏见响应,不同任务与模型间表现差异显著。据此,我们整理出更具挑战性的 FairMT-1K(1,000 条)数据集,并对 15 个前沿模型进行测试,揭示当前模型在公平性上的不足,凸显该方法在真实对话场景中评估公平性的价值,呼吁未来研究聚焦于提升模型公平性并采用 FairMT-1K 进行评估。

原文摘要 · Abstract (English)

The growing use of large language model (LLM)-based chatbots has raised concerns about fairness. Fairness issues in LLMs can lead to severe consequences, such as bias amplification, discrimination, and harm to marginalized communities. While existing fairness benchmarks mainly focus on single-turn dialogues, multi-turn scenarios, which in fact better reflect real-world conversations, present greater challenges due to conversational complexity and potential bias accumulation. In this paper, we propose a comprehensive fairness benchmark for LLMs in multi-turn dialogue scenarios, \textbf{FairMT-Bench}. Specifically, we formulate a task taxonomy targeting LLM fairness capabilities across three stages: context understanding, user interaction, and instruction trade-offs, with each stage comprising two tasks. To ensure coverage of diverse bias types and attributes, we draw from existing fairness datasets and employ our template to construct a multi-turn dialogue dataset, \texttt{FairMT-10K}. For evaluation, GPT-4 is applied, alongside bias classifiers including Llama-Guard-3 and human validation to ensure robustness. Experiments and analyses on \texttt{FairMT-10K} reveal that in multi-turn dialogue scenarios, current LLMs are more likely to generate biased responses, and there is significant variation in performance across different tasks and models. Based on this, we curate a challenging dataset, \texttt{FairMT-1K}, and test 15 current state-of-the-art (SOTA) LLMs on this dataset. The results show the current state of fairness in LLMs and showcase the utility of this novel approach for assessing fairness in more realistic multi-turn dialogue contexts, calling for future work to focus on LLM fairness improvement and the adoption of \texttt{FairMT-1K} in such efforts.

公平性评测多轮对话大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。