现有可解释方法无法可靠纠正语言模型的误判,尽管其内部知识准确率极高。
Interpretability without actionability: mechanistic methods cannot correct language model errors despite near-perfect internal representations
- 用四种机制解释方法尝试修正模型错误,但均无法稳定提升输出准确性
- 模型内部识别准确率达98.2% AUROC,但实际输出敏感度仅45.1%,差距巨大
- 方法要么无效,要么破坏正确判断,不适合用于AI安全中的错误修复
语言模型在内部表示中编码了远超其输出表现的任务相关知识,但机制可解释性方法能否弥合这一知识-行动差距尚未系统验证。我们对比了四种机制可解释性方法——概念瓶颈调控(Steerling-8B)、稀疏自编码器特征调控、对数几率透镜结合激活补丁法,以及线性探测与真实性分离向量调控(Qwen 2.5 7B Instruct)——在400个医师裁定的临床案例(144种风险,256种良性)上修正假阴性分诊错误的效果。线性探测对危险与良性案例的区分达到98.2% AUROC,但模型输出敏感度仅为45.1%,存在53个百分点的知识-行动差距。概念瓶颈调控仅纠正20%的漏报,却破坏53%的正确判断,效果与随机扰动无异(p=0.84)。SAE特征调控虽有3,695个显著特征,但未产生任何效果。高强TSV调控纠正24%漏报,仅破坏6%正确判断,但仍有76%错误未被修正。当前机制可解释方法无法可靠将内部知识转化为有效输出修正,这对依赖可解释性实现错误修正的AI安全框架提出挑战。
原文摘要 · Abstract (English)
Language models encode task-relevant knowledge in internal representations that far exceeds their output performance, but whether mechanistic interpretability methods can bridge this knowledge-action gap has not been systematically tested. We compared four mechanistic interpretability methods -- concept bottleneck steering (Steerling-8B), sparse autoencoder feature steering, logit lens with activation patching, and linear probing with truthfulness separator vector steering (Qwen 2.5 7B Instruct) -- for correcting false-negative triage errors using 400 physician-adjudicated clinical vignettes (144 hazards, 256 benign). Linear probes discriminated hazardous from benign cases with 98.2% AUROC, yet the model's output sensitivity was only 45.1%, a 53-percentage-point knowledge-action gap. Concept bottleneck steering corrected 20% of missed hazards but disrupted 53% of correct detections, indistinguishable from random perturbation (p=0.84). SAE feature steering produced zero effect despite 3,695 significant features. TSV steering at high strength corrected 24% of missed hazards while disrupting 6% of correct detections, but left 76% of errors uncorrected. Current mechanistic interpretability methods cannot reliably translate internal knowledge into corrected outputs, with implications for AI safety frameworks that assume interpretability enables effective error correction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。