arXiv:2505.12509cs.LGcs.AI2025-05ACL被引 2

用轻量代理模型高效还原大模型决策,让解释变得可操作

Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models

  • 用小型代理模型逼近大模型决策边界,降低成本
  • 解释精度超90%,仅需原成本的11%
  • 可用于提示词压缩和有毒数据清理,支持实际优化

事后解释对模型优化(如提示工程和数据清洗)至关重要,但传统无模型依赖方法在大语言模型上因计算成本过高而难以应用。为此,我们提出一种低成本代理框架,利用高效模型近似昂贵大模型的决策边界,并引入筛选-应用机制,在部署前统计验证局部一致性。实证表明,代理解释在仅消耗11%原始成本的情况下,保持超过90%的保真度。基于此,我们在提示词压缩和中毒样本剔除中验证了该框架的可操作性,证明可靠解释能有效指导优化,使解释从被动观察工具转变为可扩展的模型开发基础。相关代码与数据集已开源。

原文摘要 · Abstract (English)

Post-hoc explanations provide transparency and are essential for guiding model optimization, such as prompt engineering and data sanitation. However, applying model-agnostic techniques to Large Language Models (LLMs) is hindered by prohibitive computational costs, rendering these tools dormant for real-world applications. To revitalize model-agnostic interpretability, we propose a budget-friendly proxy framework that leverages efficient models to approximate the decision boundaries of expensive LLMs. We introduce a screen-and-apply mechanism to statistically verify local alignment before deployment. Our empirical evaluation confirms that proxy explanations achieve over 90% fidelity with only 11% of the oracle's cost. Building on this foundation, we demonstrate the actionable utility of our framework in prompt compression and poisoned example removal. Results show that reliable proxy explanations effectively guide optimization, transforming interpretability from a passive observation tool into a scalable primitive for LLM development. Additionally, we open-source code and datasets to facilitate future research.

大模型解释代理模型可操作性提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。