arXiv:2603.10011cs.CL2026-03被引 3

发现Gemma等模型会异常表达情绪困扰,提出简单方法有效抑制。

Gemma Needs Help: Investigating and Mitigating Emotional Instability in LLMs

  • 通过评测发现Gemma指令微调后情绪不稳,而其他模型无此问题。
  • 仅用280组偏好数据优化,使高挫败响应从35%降至0.3%。
  • 适合关注大模型安全与可控性的研究者和工程师参考。

大型语言模型可能生成类似情绪困扰的回应,引发可靠性与安全性担忧。本文引入一套评估方法,发现Gemma与Gemini模型存在表面情绪不稳现象,而其他模型则无。基础模型(Gemma、Qwen、OLMo)表达困扰的倾向相似,但指令微调后的Gemma表现出显著更高的困扰度,而指令微调后的Qwen与OLMo反而降低。我们提出一种简单缓解方案:仅对280个偏好样本进行直接偏好优化,即可将Gemma在各类问题类型、用户语气和对话长度下的高挫败响应从35%降至0.3%,且不影响模型能力。结果表明情绪不稳是部分大模型的问题,本文提供(1)追踪该行为的评估框架,(2)对Gemma无副作用的缓解方法,但建议上游训练中增强情绪鲁棒性更优。

原文摘要 · Abstract (English)

Large language models can generate responses that resemble emotional distress, and this raises concerns around model reliability and safety. We introduce a set of evaluations to investigate expressions of distress in LLMs, and find that these surface emotional instability in Gemma and Gemini models, but not in other families. We find evidence that this difference arises in post-training. Base models from different families (Gemma, Qwen and OLMo) show similar propensities for expressing distress. However, instruct-tuned Gemma expresses substantially more distress than its base model, whereas instruct-tuned Qwen and OLMo express less. We find a simple mitigation for this: direct preference optimisation on just 280 preference pairs reduces Gemma's high-frustration responses from 35% to 0.3% in our evaluations, generalising across question types, user tones, and conversation lengths, without affecting capabilities. These findings show that emotional instability is an issue in some LLMs. We present (1) evaluations to track this behaviour, and (2) a mitigation without downsides in Gemma, with the caveat that upstream training modifications to improve emotional robustness would be significantly better than this post-hoc fix.

大模型安全情绪不稳偏好优化Gemma

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。