arXiv:2602.14869cs.AIstat.ML2026-02被引 4

用可解释语义方向提升数据溯源效率,让模型行为更可控。

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

  • 以语义方向替代单个样本,实现更抽象的模型行为溯源
  • 新方法比经典影响函数快10倍以上,性能相当
  • 适合关注模型可解释性与训练数据控制的研究者

随着大语言模型的训练和微调日益普遍,从业者亟需识别哪些训练数据引发了特定行为,尤其是意外行为。训练数据溯源(TDA)方法通过估计数据点的影响来解决此问题。现有方法如影响函数计算成本高,且基于单个测试样本,易偏向句法而非语义相似性。为此,我们利用模型内部可解释结构进行溯源:首先提出概念影响(Concept Influence),将模型行为归因于语义方向(如线性探测器或稀疏自编码器特征),而非单个测试样本;其次证明基于探测器的方法是概念影响的一阶近似,性能相近但速度提升一个数量级以上。我们在新兴偏差基准和真实后训练数据集上验证了概念影响及其近似方法,结果表明其性能媲美经典影响函数,同时显著更高效。更广泛而言,将可解释结构融入传统TDA流程,可实现更高效、可解释且对模型行为更具控制力的数据驱动优化。

原文摘要 · Abstract (English)

As large language models are increasingly trained and fine-tuned, practitioners need methods to identify which training data drive specific behaviors, particularly unintended ones. Training Data Attribution (TDA) methods address this by estimating datapoint influence. Existing approaches like influence functions are both computationally expensive and attribute based on single test examples, which can bias results toward syntactic rather than semantic similarity. To address these issues of scalability and influence to abstract behavior, we leverage interpretable structures within the model during the attribution. First, we introduce Concept Influence which attribute model behavior to semantic directions (such as linear probes or sparse autoencoder features) rather than individual test examples. Second, we show that simple probe-based attribution methods are first-order approximations of Concept Influence that achieve comparable performance while being over an order-of-magnitude faster. We empirically validate Concept Influence and approximations across emergent misalignment benchmarks and real post-training datasets, and demonstrate they achieve comparable performance to classical influence functions while being substantially more scalable. More broadly, we show that incorporating interpretable structure within traditional TDA pipelines can enable more scalable, explainable, and better control of model behavior through data.

数据溯源可解释性大模型训练语义方向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。