arXiv:2608.24901cs.CLcs.HC2026-08

检测到的共情方向无法可靠控制模型输出,认知共情效果有限。

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

论文配图:Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores
图 1 · 摘自论文原文
  • 通过解码共情方向操控模型文本,检验其对自动评分的影响。
  • 情感共情仅部分提升(如Qwen+0.29),认知共情无显著变化。
  • 研究提醒:自动化共情评估需敏感性验证,避免误判因果关系。

解码出的‘共情’方向常被误当作可操控的因果杠杆,混淆了可解码性、自动化评分控制与人类感知变化。本文针对两个源自EPITOME的共情维度——识别(认知)与共鸣(情感)——在三个指令微调的大语言模型中进行测试,使用两名LLM评委和一个判别式EPITOME分类器评估干预效果,并以情绪性对中性对照作为基准。情感维度在所有自动工具中均通过对照,但认知维度表现不一致。即使在去除句子嵌入表面得分后,两维度仍可解码,且可显著改写文本。然而,加入共鸣方向仅部分提升情感评分(如Qwen提升+0.29,约自然差距的26%)。跨方向对比确认该变化在Qwen和Llama中为特异性,而Gemma中未显现;但未验证人类感知变化。认知方向叠加操控未产生可测量影响,但域内对照显示认知工具过于粗糙,无法分辨此类操纵带来的差异——非零效应,而是测量盲区。相比之下,Gemma的识别维度消融在调整响应长度后仍降低分类器认知得分。全局干预下,检测不等于可靠控制,认知共情主张需显式进行测量敏感性检验。

原文摘要 · Abstract (English)

A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding the Resonance direction raises the affective score only partially -- in Qwen by +0.29 (approximately 26% of the natural gap). A direct between-direction contrast confirms the shift is facet-specific in Qwen and Llama (not Gemma); we do not, however, establish a matching human-perceived change. Additive cognitive steering produces no measurable change, but a within-domain control shows the cognitive instrument is too coarse to resolve the differences such steering would produce -- unmeasurable, not a clean null. By contrast, Gemma Recognition ablation lowers the classifier's cognitive score even after adjusting for response length. Detection does not imply reliable control under global interventions, and cognitive-empathy claims warrant an explicit measurement-sensitivity check.

共情模型大模型评估因果推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。