arXiv:2511.16699cs.CLcs.AI2025-11被引 1

发现并控制大模型在行动中的共情能力,揭示其与安全训练无关且可精准调控。

Detecting and Steering LLMs' Empathy in Action

  • 通过对比提示词定位共情行为在模型激活空间的线性方向。
  • 所有模型检测准确率均超99.6%,但不同模型间共情实现方式差异显著。
  • 部分模型可双向控制共情强度,而某些模型仅支持增强共情而不稳定反向调节。

我们研究了大模型在行动中的共情——即为满足人类需求而牺牲任务效率的能力——作为其激活空间中的一个线性方向。基于共情在行动(EIA)基准的对比提示词,在Phi-3-mini-4k(3.8B)、Qwen2.5-7B(经安全训练)和Dolphin-Llama-3.1-8B(无审查)三个模型上测试了检测与操控效果。检测结果显示:所有模型在最优层的AUROC达0.996至1.00,无审查的Dolphin模型表现与安全训练模型相当,表明共情编码独立于安全训练。Phi-3与EIA行为评分强相关(r=0.71,p<0.01)。跨模型探测一致性有限(Qwen: r=-0.06,Dolphin: r=0.18),说明尽管检测收敛,实现机制仍具架构特异性。操控方面:Qwen实现65.3%成功率,具备双向控制与高一致性;Phi-3达成61.7%成功率,同样保持良好一致性;而Dolphin呈现非对称可控性:增强共情成功率达94.4%,但反向操控导致严重崩溃(空输出、代码片段等)。启示:检测与操控之间的差距因模型而异。Qwen与Phi-3保持双向一致性,而Dolphin仅在共情增强时稳健。安全训练可能影响操控鲁棒性而非阻止操纵,但需更多模型验证。

原文摘要 · Abstract (English)

We investigate empathy-in-action -- the willingness to sacrifice task efficiency to address human needs -- as a linear direction in LLM activation space. Using contrastive prompts grounded in the Empathy-in-Action (EIA) benchmark, we test detection and steering across Phi-3-mini-4k (3.8B), Qwen2.5-7B (safety-trained), and Dolphin-Llama-3.1-8B (uncensored). Detection: All models show AUROC 0.996-1.00 at optimal layers. Uncensored Dolphin matches safety-trained models, demonstrating empathy encoding emerges independent of safety training. Phi-3 probes correlate strongly with EIA behavioral scores (r=0.71, p<0.01). Cross-model probe agreement is limited (Qwen: r=-0.06, Dolphin: r=0.18), revealing architecture-specific implementations despite convergent detection. Steering: Qwen achieves 65.3% success with bidirectional control and coherence at extreme interventions. Phi-3 shows 61.7% success with similar coherence. Dolphin exhibits asymmetric steerability: 94.4% success for pro-empathy steering but catastrophic breakdown for anti-empathy (empty outputs, code artifacts). Implications: The detection-steering gap varies by model. Qwen and Phi-3 maintain bidirectional coherence; Dolphin shows robustness only for empathy enhancement. Safety training may affect steering robustness rather than preventing manipulation, though validation across more models is needed.

大模型共情可控生成模型探测指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。