arXiv:2602.17867cs.LGcs.CL2026-02

提出ADAPT方法,让大模型特征可视化更高效准确

ADAPT: Hybrid Prompt Optimization for LLM Feature Visualization

  • 融合束搜索与自适应梯度突变,克服文本优化易陷局部最优问题
  • 在Gemma 2 2B的稀疏自编码器潜空间中,性能全面优于现有方法
  • 适用于研究模型内部表征,尤其适合关注特征可解释性的研究者

理解大语言模型激活空间中学习方向所编码的特征,需要找到能强烈激活它们的输入。特征可视化通过优化输入以最大化激活目标方向,为替代昂贵的数据集搜索提供了新路径,但受限于文本的离散特性,在大模型中仍处于探索阶段。现有提示优化技术在此领域表现不佳,极易陷入局部极小值。为此,我们提出ADAPT,一种结合束搜索初始化与自适应梯度引导突变的混合方法,针对该领域的失效模式进行设计。我们在Gemma 2 2B的稀疏自编码器潜空间上进行了评估,提出了基于数据集激活统计的评价指标,实现严谨对比,结果表明ADAPT在不同层和潜变量类型上均持续优于先前方法。研究证明,大模型的特征可视化是可行的,但需采用符合领域特性的设计假设。

原文摘要 · Abstract (English)

Understanding what features are encoded by learned directions in LLM activation space requires identifying inputs that strongly activate them. Feature visualization, which optimizes inputs to maximally activate a target direction, offers an alternative to costly dataset search approaches, but remains underexplored for LLMs due to the discrete nature of text. Furthermore, existing prompt optimization techniques are poorly suited to this domain, which is highly prone to local minima. To overcome these limitations, we introduce ADAPT, a hybrid method combining beam search initialization with adaptive gradient-guided mutation, designed around these failure modes. We evaluate on Sparse Autoencoder latents from Gemma 2 2B, proposing metrics grounded in dataset activation statistics to enable rigorous comparison, and show that ADAPT consistently outperforms prior methods across layers and latent types. Our results establish that feature visualization for LLMs is tractable, but requires design assumptions tailored to the domain.

特征可视化大模型解释提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。