arXiv:2602.23993cs.CL2026-02

用梯度方法自动学习语言模型中的特征方向,支持训练、评估与可视化。

The GRADIEND Python Package: An End-to-End System for Gradient-Based Feature Learning

  • 基于掩码语言模型梯度挖掘特征方向
  • 支持多特征对比与可控权重重写
  • 适合需要可解释性的大模型研究者

我们提出gradiend,一个开源Python工具包,实现了从事实-反事实掩码语言模型(MLM)和上下文语言模型(CLM)梯度中学习特征方向的GRADIEND方法。该工具包提供统一工作流,涵盖特征相关数据生成、训练、评估、可视化、通过受控权重更新实现模型持久化改写以及多特征比较。通过英文代词实例、语义情感案例(评估词汇泛化到未见目标词的能力)及大规模特征对比,验证了gradiend的有效性。

原文摘要 · Abstract (English)

We present gradiend, an open-source Python package that operationalizes the GRADIEND method for learning feature directions from factual-counterfactual MLM and CLM gradients in language models. The package provides a unified workflow for feature-related data creation, training, evaluation, visualization, persistent model rewriting via controlled weight updates, and multi-feature comparison. We demonstrate gradiend through an English pronoun running example, a semantic sentiment use case that evaluates lexical generalization to held-out target words, and a large-scale feature comparison.

特征学习可解释性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。