arXiv:2601.03786cs.CLcs.LG2026-01ACL被引 1

提出新方法筛选训练数据,让模型解释更精准。

Compact Example-Based Explanations for Language Models

  • 设计无需重训练的选例相关性评分
  • 实验证明常见策略不如随机选择
  • 兼顾影响度与代表性,提升解释效果

训练数据影响估计方法可量化训练文档对模型输出的贡献,是生成基于示例解释的潜在来源。由于人类无法理解数千篇文档,解释中只能呈现少量训练数据。尽管选择哪些文档直接影响解释质量,但以往评估未关注选择策略。为此,我们提出一种新的选择相关性评分,这是一种无需重训练的度量,用于衡量一组示例在解释模型输出时的有用性。通过微调实验验证,该评分能准确预测示例集是否支持或削弱模型预测。利用此评分,我们发现常见选择策略往往表现劣于随机选择。基于这一发现,我们提出一种平衡影响度与代表性的策略,相比直接选取排名最高者,能更高效利用有限的示例预算。

原文摘要 · Abstract (English)

Training data influence estimation methods quantify the contribution of training documents to a model's output, making them a promising source of information for example-based explanations. As humans cannot interpret thousands of documents, only a small subset of the training data can be presented as an explanation. Although the choice of which documents to include directly affects explanation quality, previous evaluations of such systems have largely ignored any selection strategies. To address this, we propose a novel selection relevance score, a retraining-free metric that quantifies how useful a set of examples is for explaining a model's output. We validate this score through fine-tuning experiments, confirming that it can predict whether a set of examples supports or undermines the model's predictions. Using this metric, we further show that common selection strategies often underperform random selection. Motivated by this finding, we propose a strategy that balances influence and representativeness, enabling better use of selection budgets than naively selecting the highest-ranking examples.

模型解释示例选择影响度估计文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。