arXiv:2606.08236cs.CLcs.LG2026-06中稿 · ICML

不依赖目标输出,自动发现模型续写背后的语义与机制模式。

Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms

论文配图:Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms
图 1 · 摘自论文原文
  • 用语义嵌入和机制归因联合表征续写内容,无监督聚类。
  • 在多个数据集上识别出单视角方法遗漏的续写模式。
  • 适合做模型机制审计,尤其关注高风险场景下的行为一致性。

随着大语言模型在高风险场景中广泛应用,亟需能够审计模型输出及其内部计算过程的工具。电路分析是机械可解释性中的核心方法,但通常针对特定提示-完成对进行解释,难以揭示模型续写分布的异质性。本文提出分布级无监督特征发现方法,通过同时利用语义内容与序列级机制归因,对采样得到的续写进行聚类,无需人工指定目标输出。该方法为每个续写表示为语义嵌入和前缀到续写的归因签名,并优化一个率失真目标,权衡语义连贯性、机制一致性与聚类粒度。在聚类与干预分析中,所发现的簇揭示了单视图基线遗漏的续写模式,并提供了证据表明簇签名对应可操作的机制因素。整体而言,该方法补充了电路分析与行为评估,为模型续写分布背后的机制提供了可扩展的审计能力。

原文摘要 · Abstract (English)

As large language models are increasingly deployed in high-stakes settings, there is a growing need for tools that audit not only model outputs but also the internal computations that produce them. Circuit analysis is a central approach in mechanistic interpretability, but it is typically target-conditioned, explaining a single prompt paired with a chosen completion. This target-conditioned setup can obscure heterogeneity across a model's continuation distribution. We introduce distribution-level unsupervised feature discovery, which clusters sampled continuations using both semantic content and sequence-level mechanistic attributions, without manually specifying target outputs. Our method represents each continuation with a semantic embedding and a prefix-to-continuation attribution signature, then optimizes a rate-distortion objective that trades off semantic coherence, mechanistic consistency, and cluster granularity. Across clustering and steering analyses, the discovered clusters expose continuation modes that single-view baselines miss and provide interventional evidence that cluster signatures correspond to actionable mechanistic factors. Overall, our approach complements circuit analysis and behavioral evaluation by providing a scalable audit of the mechanisms underlying a model's continuation distribution.

可解释性无监督学习机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。