揭示对话中隐蔽的善意偏见,发现现有检测工具难以识别伪装成友好的不平等对待。
Benevolent Bias in Multi-Turn Human-Agent Dialogue

- 定义善意偏见:表面友好但实际不平等的对话行为
- 构建36万条多轮对话数据集,覆盖多种用户与代理特征
- 现有工具易漏判善意偏见,大模型判别器更易误判中性支持
人类-代理对话中的偏见不仅表现为攻击性语言,也可能以善意偏见形式出现,即在温暖积极的语气下隐藏不平等对待。为此,我们从语气和对待方式两个维度操作化定义善意偏见,形成三类:中立支持、明显偏见和善意偏见。基于此,我们构建了BENEVDIAL数据集,包含362,880条多轮支持性对话,涵盖用户与代理的人口统计学特征、角色及生成器类型,支持受控评估。在此基础上,我们测试了两类检测器:现成的安全检测工具和提示式大语言模型(LLM)判别器。结果表明存在检测差距:现成检测工具能可靠识别明显偏见,但基本忽略善意偏见;而LLM判别器在更明确的标准下可捕捉更多偏见,却逐渐将中立支持误判为善意偏见,且人口统计背景会加剧误报。这些发现表明,公平的对话监控必须超越表层线索,关注代理的实际对待是否存在差异。
原文摘要 · Abstract (English)
Bias in human-agent interaction can manifest not only through hostile language but also as benevolent bias, whereby unequal treatment hides behind a warm, positive tone. To make it detectable, we operationalise benevolent bias along two dimensions, tone and treatment, yielding three classes: neutral support, overt bias, and benevolent bias. Building on these definitions, we construct BENEVDIAL, a class-balanced corpus of 362,880 multi-turn support dialogues spanning user and agent demographics, roles, and generators, to support controlled evaluation. We then test two detector families on it: off-the-shelf safety detectors and prompted large language model (LLM) judges. Our results reveal a detection gap: off-the-shelf detectors reliably flag overt bias yet largely miss benevolent bias, while LLM judges catch more under more explicit detection criteria but increasingly misclassify neutral support as benevolent bias, and demographic context amplifies the false alarms. These findings suggest that fair monitoring of human-agent dialogue must look beyond surface cues to whether the agent's treatment is disparate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。