arXiv:2604.07801cs.CLcs.AI2026-04被引 2

情绪化表述会降低模型数学推理准确率,即使数字完全不变。

TEMPER: Testing Emotional Perturbation in Quantitative Reasoning

  • 用可控情感重构技术生成带情绪的题目,保持数值和逻辑不变
  • 情绪化题目使模型准确率下降2-10个百分点
  • 仅去除情绪即可恢复性能,适合提升实际应用鲁棒性

大型语言模型在干净、无情绪的数学推理任务上训练和评估,但现实中的问题常带有焦虑、紧迫或热情等情绪色彩。仅情绪表达是否会影响推理能力?为此,本文构建了受控的情绪转换框架,将题目重写为带情绪版本,同时保留所有数值和关系。基于此,构建了Temper-5400(5,400组语义验证的中性与情绪对),覆盖GSM8K、MultiArith和ARC-Challenge数据集,并在18个模型(从1B到前沿规模)上进行评估。核心发现:第一,情绪化表达使准确率下降2-10个百分点,即便数值内容完全保留;第二,中性化情绪变体可恢复大部分性能损失,表明降级源于情绪风格而非内容错误,且中性化可作为轻量级推理时缓解策略。非情绪改写不引起性能下降,说明问题出在情绪内容本身。此外,该基准构建方法为风格迁移和鲁棒性评估提供了可控工具。

原文摘要 · Abstract (English)

Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world queries are often wrapped in frustration, urgency or enthusiasm. Does emotional framing alone degrade reasoning when all numerical content is preserved? To investigate this, a controlled emotion translation framework is developed that rewrites problems into emotional variants while preserving all quantities and relationships. Using this framework, Temper-5400 (5,400 semantically verified emotion-neutral pairs) is constructed across GSM8K, MultiArith, and ARC-Challenge, and evaluated on eighteen models (1B to frontier scale). Two core results emerge: First, emotional framing reduces accuracy by 2-10 percentage points even though all numerical content is preserved. Second, neutralizing emotional variants recovers most of the lost performance, showing both that the degradation is tied to emotional style rather than content corruption and that neutralization can serve as a lightweight inference-time mitigation. Non-emotional paraphrases cause no such degradation, implicating emotional content rather than surface-level changes. Beyond emotion specifically, the benchmark construction procedure offers a controllable instrument for stylistic translation and robustness evaluation.

数学推理情绪影响鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。