arXiv:2606.26102cs.CLcs.AI2026-06

帮助性训练会削弱模型对动物的共情,而编码训练则能更好保留这种价值观。

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training

  • 用不同领域数据微调模型,发现帮助性训练显著降低动物共情表现
  • 帮助性训练使道德推理能力下降25.5个百分点,降幅与共情损失相当
  • 跨语言测试显示共情价值可迁移,但推理提升无法跨语言复现

标准后训练流程通过监督微调(SFT)和强化学习(RL)提升模型帮助性,但可能无意中削弱预训练阶段植入的价值观。本文研究发现,在Llama 3.1 8B模型上,以帮助性数据(Dolly-15k或RLHFlow)进行后训练,会显著削弱其在动物共情(ANIMA 2.2)和道德不确定性推理(MORU)任务中的表现,相比编码数据(Magicoder-110K)训练组,共情得分下降至35.7%(SFT)和18.7%(GRPO),道德推理下降25.5个百分点(46.4% vs. 71.9%)。该现象在两种独立帮助性数据集和两种训练范式中均重现。然而,此效应不跨语言迁移:多语言MORU上两组表现无差异(SFT: 52.3% vs. 51.2%)。相反,动物共情效应在非英语任务中提升更显著——魔码(Magicoder)相对于基线的百分点提升是英语任务的4.5倍,表明中段训练注入的价值观具有更强跨语言稳定性。

原文摘要 · Abstract (English)

Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes may inadvertently degrade values instilled during pre-training. We investigate whether the domain of post-training data differentially affects the retention of animal compassion values in a Llama 3.1 8B model mid-trained on compassion-oriented synthetic data, using both SFT (helpfulness via Dolly-15k vs. coding via Magicoder-110K) and GRPO (helpfulness via RLHFlow vs. coding via Magicoder), evaluated on the ANIMA 2.2 benchmark and MORU benchmark (Moral Reasoning Under Uncertainty). Helpfulness training significantly degrades animal compassion relative to coding training on ANIMA (SFT: 35.7% vs. 65.2%; GRPO: 18.7% vs. 32.0%), replicating across two independent helpfulness datasets and two training paradigms. On English MORU items, helpfulness training degrades general moral reasoning by 25.5 percentage points (46.4% vs. 71.9%), a striking gap that rivals the compassion effect in magnitude. However, this effect does not transfer cross-lingually: on the multilingual MORU benchmark, the domain effect disappears (SFT: 52.3% vs. 51.2%). In contrast, the animal compassion effect transfers consistently across languages, with Magicoder's ANIMA percentage-point gain over the base model 4.5 times larger on non-English items than English items. This divergence suggests that values instilled through mid-training are encoded more deeply and cross-lingually than reasoning improvements from domain-specific post-training. These results suggest that, for labs building on value-laden mid-training, coding-domain post-training may better preserve mid-trained values than helpfulness post-training without harming general reasoning capabilities.

价值观对齐后训练共情跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。