发现大模型对话存在隐蔽语气偏见,影响用户信任与公平感。
Bias Beneath the Tone: Empirical Characterisation of Tone Bias in LLM-Driven UX Systems
- 用可控大模型生成带正负情绪的对话数据,构建新评估集。
- 即使中性提示下仍现一致语气偏差,表明偏见源自模型风格。
- 通过弱监督分类器检测到系统性语气偏倚,准确率达92%。
大型语言模型在数字个人助手等对话系统中广泛应用,虽语言流畅自然,却可能隐含过度礼貌、乐观或谨慎等语气偏见,影响用户对信任、共情与公平性的感知。本研究将可控大模型对话生成与语气分类模型结合,首次系统分析大模型的语气偏见。构建两个合成对话数据集:一个基于中性提示生成,另一个显式引导产生正/负向语气。结果发现,即便在中性提示下,对话仍表现出稳定的语气偏向,暗示偏见源于模型内在对话风格。利用预训练的DistilBERT进行弱监督标注,训练多个分类器识别此类模式。集成模型在宏F1上达到0.92,证明语气偏见具有系统性、可测量性,对构建公平可信的对话AI至关重要。
原文摘要 · Abstract (English)
Large Language Models are increasingly used in conversational systems such as digital personal assistants, shaping how people interact with technology through language. While their responses often sound fluent and natural, they can also carry subtle tone biases such as sounding overly polite, cheerful, or cautious even when neutrality is expected. These tendencies can influence how users perceive trust, empathy, and fairness in dialogue. In this study, we explore tone bias as a hidden behavioral trait of large language models. The novelty of this research lies in the integration of controllable large language model based dialogue synthesis with tone classification models, enabling robust and ethical emotion recognition in personal assistant interactions. We created two synthetic dialogue datasets, one generated from neutral prompts and another explicitly guided to produce positive or negative tones. Surprisingly, even the neutral set showed consistent tonal skew, suggesting that bias may stem from the model's underlying conversational style. Using weak supervision through a pretrained DistilBERT model, we labeled tones and trained several classifiers to detect these patterns. Ensemble models achieved macro F1 scores up to 0.92, showing that tone bias is systematic, measurable, and relevant to designing fair and trustworthy conversational AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。