arXiv:2512.14719cs.LGcs.AI2025-12

提出新方法提升小模型分类时的可解释性与鲁棒性。

Hybrid Attribution Priors for Explainable and Robust Model Training

  • 设计细粒度类感知归因先验,聚焦关键差异特征。
  • 在全数据、少样本和对抗场景中均提升模型区分能力。
  • 适合需要高可解释性与抗干扰能力的轻量级应用。

小型语言模型(SLMs)广泛应用于低延迟、轻量部署的分类任务。随着可解释性与鲁棒性重要性提升,基于归因的监督学习成为有效框架,但通用可靠的归因先验仍难获取。分析发现,现有归因方法虽能定位类相关词元,却常关注语义相似类别共享的常见关键词,这类信息在标准训练下已难以区分,导致归因缺乏判别力。为此,本文提出类感知归因先验(CAP),引导模型捕捉细微类别差异,生成更显著、更具判别性的归因先验。进一步提出CAP Hybrid,融合CAP与现有归因方法的先验,构建更全面平衡的监督信号。通过对齐模型自归因与这些增强先验,促进多样且决策相关特征的学习。大量实验表明,该方法在全数据、少样本及对抗场景中持续提升模型可解释性与鲁棒性。

原文摘要 · Abstract (English)

Small language models (SLMs) are widely used in tasks that require low latency and lightweight deployment, particularly classification. As interpretability and robustness gain increasing importance, explanation-guided learning has emerged as an effective framework by introducing attribution-based supervision during training; however, deriving general and reliable attribution priors remains a significant challenge. Through an analysis of representative attribution methods in classification settings, we find that although these methods can reliably highlight class-relevant tokens, they often focus on common keywords shared by semantically similar classes. Because such classes are already difficult to distinguish under standard training, these attributions provide insufficient discriminative cues, limiting their ability to improve model differentiation. To overcome this limitation, we propose Class-Aware Attribution Prior (CAP), a novel attribution prior extraction framework that guides language models toward capturing fine-grained class distinctions and producing more salient, discriminative attribution priors. Building on this idea, we further introduce CAP Hybrid, which combines priors from CAP with those from existing attribution techniques to form a more comprehensive and balanced supervisory signal. By aligning a model's self-attribution with these enriched priors, our approach encourages the learning of diverse, decision-relevant features. Extensive experiments in full-data, few-shot, and adversarial scenarios demonstrate that our method consistently enhances both interpretability and robustness.

可解释性小模型归因鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。