arXiv:2510.27413cs.LGcs.AI2025-10被引 2

用已有概念图谱让新模型自动可解释,省去重复标注成本。

Atlas-Alignment: Making Interpretability Transferable Across Language Models

  • 通过轻量对齐将新模型隐空间映射到预训练概念图谱。
  • 无需标签数据即可实现语义检索与可控生成,效果接近专用模型。
  • 适合快速部署可解释模型的研究者和工业应用团队。

可解释性对构建安全、可靠、可控的语言模型至关重要,但现有解释流程成本高且难以扩展。解释新模型通常需训练特定组件(如稀疏自编码器),再经人工或半自动标注验证,带来持续增长的“透明度税”,无法跟上模型发展速度。我们提出Atlas-Alignment框架,通过仅使用共享输入和轻量级表征对齐方法,将新模型的隐空间对齐至预先构建并标注好的概念图谱(Concept Atlas),避免重复投入。定量与定性评估表明,简单对齐方法即可实现稳健的语义检索和可调控生成,无需额外标注概念数据集。因此,Atlas-Alignment实现了可解释AI与机制可解释性的成本分摊:只需投资一次高质量概念图谱,即可以极低边际成本使多个新模型具备可解释性与可控性。

原文摘要 · Abstract (English)

Interpretability is crucial for building safe, reliable, and controllable language models, yet existing interpretability pipelines remain costly and difficult to scale. Interpreting a new model typically requires training model-specific components (e.g., sparse autoencoders), followed by manual or semi-automated labeling and validation, imposing a growing "transparency tax" that does not scale with the pace of model development. We introduce Atlas-Alignment, a framework that avoids this cost by aligning the latent space of a new model to a pre-existing, labeled Concept Atlas using only shared inputs and lightweight representational alignment methods. Through quantitative and qualitative evaluations, we show that simple alignment methods enable robust semantic retrieval and steerable generation without the need for labeled concept datasets. Atlas-Alignment thus amortizes the cost of explainable AI and mechanistic interpretability: by investing in a single high-quality Concept Atlas, we can make many new models transparent and controllable at minimal marginal cost.

可解释性模型对齐概念图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。