让情绪解释真正对应真实因果,而非只是看起来合理。
Faithful Action-unit Causal Reasoning for Counterfactually Faithful Emotion Explanations
- 基于因果图设计反事实训练目标,强制模型只引用真正影响情绪的面部动作单元。
- 在疼痛表情数据集上,解释一致性从0.08提升至0.57,跨数据集仍保持高一致性。
- 可生成忠实的自然语言解释,且适用于多语言模型,适合可信AI研究者使用。
多模态模型能识别引发面部情绪的动作单元(AUs),但其解释往往看似合理却未必真实:模型引用的AUs未必是实际驱动预测的原因。本文将AU→情绪推理视为理由、标签与结构化因果图G之间的反事实一致性问题,提出FACR方法,通过独立推导的极性感知因果图G来约束解释器,并训练反事实忠实性目标:对图中标记为因果的AU施加do干预应改变预测,标记为无关的则不应改变。忠实性由此可训练且可度量,我们以已知因果结构PSPI疼痛-AU构成为基准进行评估,现有情感推理基准无法支持此测试。该指标检验模型是否遵循给定结构,而非重新发现它:在未见样本和另一数据集上,评估模型是否引用了结构所标记的因果AU。在UNBC-PAIN上进行主体无关评估,引入该目标后,模型引用的AUs与PSPI构成交集从基线0.08提升至0.57,检测代价微小;无忠实事例控制组证明该提升归因于该目标。在跨数据集情绪迁移任务中,七分类任务下对图的忠实性从0.50提升至0.84。最后,结合语言生成器,通过根据潜在激活值调节每个AU的输出,使解释天然忠实:删除某个AU即从解释中消失,该特性可迁移到第二语言模型,而自由生成的解释则不忠实。
原文摘要 · Abstract (English)
Multimodal models can name the action units (AUs) behind a facial emotion, but their AU->emotion rationales are typically plausible rather than faithful: nothing forces the AUs a model invokes to be the AUs that actually drive its prediction. We cast AU->emotion reasoning as a counterfactual-consistency problem between the rationale, the label, and a structural AU->emotion causal graph G, and propose FACR, which grounds the reasoner in an independently induced, polarity-aware G and trains a counterfactual-faithfulness objective: a do-intervention on an AU that G marks causal for a class must move the prediction, while one it marks irrelevant must leave it unchanged. Faithfulness is thereby both trainable and measurable through a matching interventional metric, which we evaluate against a known causal structure, the PSPI pain-AU composition, as no existing affective-reasoning benchmark allows. We are explicit that this metric tests fidelity to the supplied structure rather than its rediscovery: it asks whether the trained reasoner invokes the AUs the structure marks causal, on held-out subjects and a second dataset. Under subject-independent evaluation on UNBC-PAIN, the objective raises the agreement between the invoked AUs and the PSPI composition from a no-objective baseline of 0.08 to 0.57, at a small detection cost; an unfaithfulness control attributes the gain to the objective. On a cross-dataset emotion transfer, the objective likewise raises fidelity to G on a seven-class task (0.50 to 0.84). Finally, we attach a language verbalizer and extend the audit to the generated text: biasing each action unit's emission by its latent activation makes the rationale faithful by construction, so that ablating an AU removes it from the explanation, a property that transfers to a second language-model backbone, whereas a freely generated rationale is unfaithful.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。