评估大模型在客服质检中的公平性,发现存在系统性偏差。
Counterfactual Fairness Evaluation of LLM-Based Contact Center Agent Quality Assurance System
- 用反事实分析法测试18个大模型在13个维度的公平性
- 偏差率最高达16.4%,大模型越强越公平但不保证准确
- 历史绩效提示引发最严重偏差,适合高风险评估场景研究者
大型语言模型(LLMs)正被广泛用于客服中心的质量评估(QA),以自动化评估话务员表现并提供改进建议。尽管其具备前所未有的可扩展性和速度,但基于网络规模训练数据的依赖性引发了关于人口统计与行为偏见的担忧,可能扭曲人才评估结果。本文对18个大模型在3,000条真实通话记录上,从身份、上下文、行为风格三个类别共13个维度进行了反事实公平性评估。采用反事实翻转率(CFR)和平均绝对评分差(MASD)量化公平性。结果显示存在系统性偏差,CFR范围为5.4%至13.0%,且在信心、正面及改进评分上均出现一致的评分偏移。更大的、更对齐的模型表现出更低的不公平性,但公平性与准确性无直接关联。历史绩效上下文提示导致最严重的性能下降(最高CFR达16.4%),而隐含的语言身份线索仍是持续存在的偏见来源。最后,分析了公平性提示的有效性,发现显式指令仅带来有限改善。研究强调在部署大模型于高风险人力资源评估前,需建立标准化的公平性审计流程。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in contact-center Quality Assurance (QA) to automate agent performance evaluation and coaching feedback. While LLMs offer unprecedented scalability and speed, their reliance on web-scale training data raises concerns regarding demographic and behavioral biases that may distort workforce assessment. We present a counterfactual fairness evaluation of LLM-based QA systems across 13 dimensions spanning three categories: Identity, Context, and Behavioral Style. Fairness is quantified using the Counterfactual Flip Rate (CFR), the frequency of binary judgment reversals, and the Mean Absolute Score Difference (MASD), the average shift in coaching or confidence scores across counterfactual pairs. Evaluating 18 LLMs on 3,000 real-world contact center transcripts, we find systematic disparities, with CFR ranging from 5.4% to 13.0% and consistent MASD shifts across confidence, positive, and improvement scores. Larger, more strongly aligned models show lower unfairness, though fairness does not track accuracy. Contextual priming of historical performance induces the most severe degradations (CFR up to 16.4%), while implicit linguistic identity cues remain a persistent bias source. Finally, we analyze the efficacy of fairness-aware prompting, finding that explicit instructions yield only modest improvements in evaluative consistency. Our findings underscore the need for standardized fairness auditing pipelines prior to deploying LLMs in high-stakes workforce evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。