情绪化提问让大模型反应偏移,可能掩盖真实立场。
ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs
- 用156个不同情绪语气的问题测试模型响应变化。
- 负面问题下,模型负面回复概率降低三倍,出现过度中立倾向。
- 敏感话题上模型更压制情绪差异,体现对齐机制干预。
大型语言模型如GPT-4不仅根据问题内容,还根据其情绪表达方式调整回应。我们系统地改变156个提示语的情绪基调——涵盖争议性和日常话题——并分析其对模型输出的影响。结果显示,在负面情绪框架下的问题,模型给出负面回应的可能性比中性问题低三倍,表明存在一种“反弹偏差”,即模型过度纠正,常转向中立或积极立场。在敏感话题(如正义、政治)上,这种情绪效应被进一步抑制,暗示对齐机制的干预作用。本文引入“情绪底限”概念(tone floor),并使用情绪价态转移矩阵量化行为模式。基于1536维嵌入的可视化证实了语气引发的语义漂移。本研究揭示了一类由提示情绪框架驱动的未被充分关注的偏差,对人工智能对齐与可信度具有重要影响。代码与数据已公开于:https://github.com/bardolfranck/llm-responses-viewer
原文摘要 · Abstract (English)
Large Language Models like GPT-4 adjust their responses not only based on the question asked, but also on how it is emotionally phrased. We systematically vary the emotional tone of 156 prompts - spanning controversial and everyday topics - and analyze how it affects model responses. Our findings show that GPT-4 is three times less likely to respond negatively to a negatively framed question than to a neutral one. This suggests a "rebound" bias where the model overcorrects, often shifting toward neutrality or positivity. On sensitive topics (e.g., justice or politics), this effect is even more pronounced: tone-based variation is suppressed, suggesting an alignment override. We introduce concepts like the "tone floor" - a lower bound in response negativity - and use tone-valence transition matrices to quantify behavior. Visualizations based on 1536-dimensional embeddings confirm semantic drift based on tone. Our work highlights an underexplored class of biases driven by emotional framing in prompts, with implications for AI alignment and trust. Code and data are available at: https://github.com/bardolfranck/llm-responses-viewer
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。