通过可解释图与因果探测,实现多模态生成模型的机制洞察与偏见修复。
Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- 构建归属图与因果探测,定位模型内部因果结构。
- 在多个数据集上实现94.1%准确率与41%偏见降低。
- 适合关注模型可解释性、公平性与生成质量的研究者。
我们将生成模型内部视为可解析的机制对象而非黑箱,提出归属图(AGs),将GradCAM++扩展至电路级表征,并引入基于do-计算的因果探测方法,以识别因果潜在结构,实现在训练中检测与修正虚假相关、人口统计学偏见及决策电路错位。进一步提出认知对齐得分(CAS)量化模型内部表示与人类概念的一致性,设计仅共享阈值化归属节点的“先显著后隐私”机制,采用偏见感知正则化对齐子群体统计,以及融合归属信号的“揭示-修正”循环,直接整合至参数更新而无需额外微调。在CelebA、FairFace、Jigsaw和HateXplain数据集上评估,方法达到94.1%准确率、92.3%宏平均F1、79.4% IoU-XAI和12.7的FID,在72–76%对抗鲁棒性下,子群体偏见差Δ_bias降低41%,证明机制可解释性、公平性与生成性能可协同优化。
原文摘要 · Abstract (English)
We treat the internals of generative models as mechanistic objects rather than black boxes. We introduce \textbf{Attribution Graphs} (AGs), which extend GradCAM++ to circuit-level representations, and \textbf{Causal Probing}, a do-calculus intervention method for identifying causal latent structures, enabling detection and correction of spurious correlations, demographic biases, and misaligned decision circuits during training. We further propose the \textbf{Cognitive Alignment Score (CAS)}, quantifying agreement between model-internal representations and human concepts, a \textbf{saliency-first privacy mechanism} sharing only thresholded attribution nodes, a bias-aware regularizer aligning subgroup statistics, and a Reveal-to-Revise loop integrating attribution signals into parameter updates without separate fine-tuning. Evaluated on CelebA, FairFace, Jigsaw, and HateXplain, our method achieves \textbf{94.1\%} accuracy, \textbf{92.3\%} macro F1, \textbf{79.4\%} IoU-XAI, and \textbf{12.7} FID at 72--76\% adversarial robustness, while reducing subgroup disparity $Δ_{\mathrm{bias}}$ by \textbf{41\%}, demonstrating that mechanistic interpretability, fairness, and generative performance can be jointly optimized.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。