让模型解释可操作:通过激活调控实现视觉模型的精准调试
From Attribution to Action: A Human-Centered Application of Activation Steering

- 结合SAE归因与激活调控,构建可交互的实例级概念分析流程
- 8位专家均实现从查看到干预的思维转变,7人采用组件抑制策略
- 适合模型调试者、可解释性研究者,但需警惕调控的连锁效应
可解释人工智能(XAI)方法能揭示影响模型预测的特征,但难以指导实践者采取行动。基于XAI识别的组件进行激活调控,为可操作的解释提供了可能,但其实际效用尚未充分研究。本文提出一种结合SAE归因与激活调控的交互式工作流,用于视觉模型中概念使用的实例级分析,并开发为网页工具。基于该工具,我们对CLIP模型开展了半结构化专家访谈(N=8),研究从业者如何理解、信任并应用激活调控。结果显示,所有参与者均实现了从检查到干预式假设验证的转变,其中6人更依赖模型响应结果而非解释合理性来建立信任;多数人采用组件抑制的系统性调试策略,同时指出调控存在涟漪效应及实例修正泛化能力有限等风险。总体而言,激活调控使可解释性更具操作性,但也带来了安全使用的重要考量。
原文摘要 · Abstract (English)
Explainable AI (XAI) methods reveal which features influence model predictions, yet provide limited means for practitioners to act on these explanations. Activation steering of components identified via XAI offers a path toward actionable explanations, although its practical utility remains understudied. We introduce an interactive workflow combining SAE-based attribution with activation steering for instance-level analysis of concept usage in vision models, implemented as a web-based tool. Based on this workflow, we conduct semi-structured expert interviews (N=8) with debugging tasks on CLIP to investigate how practitioners reason about, trust, and apply activation steering. We find that steering enables a shift from inspection to intervention-based hypothesis testing (8/8 participants), with most grounding trust in observed model responses rather than explanation plausibility alone (6/8). Participants adopted systematic debugging strategies dominated by component suppression (7/8) and highlighted risks including ripple effects and limited generalization of instance-level corrections. Overall, activation steering renders interpretability more actionable while raising important considerations for safe and effective use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。