arXiv:2504.01771cs.AIcs.LG2025-04被引 2

通过搜索方法分析训练数据对生成结果的影响,提升模型可解释性。

Enhancing Interpretability in Generative AI Through Search-Based Data Influence Analysis

  • 用搜索思路找训练数据对生成内容的影响
  • 能定位影响输出的关键训练子集
  • 适合关注版权与艺术创作的模型使用者

生成式AI模型虽功能强大,但缺乏透明性,尤其在涉及艺术或版权内容时尤为关键。本文提出一种受搜索启发的方法,通过分析训练数据对模型输出的影响来增强可解释性。该方法聚焦于模型输出而非内部状态,同时考虑原始数据和潜在空间嵌入来搜索数据项的影响。通过局部重训练和揭示训练数据中关键影响子集,验证了方法的有效性。本工作为未来基于领域专家的用户评估等扩展奠定了基础,有望进一步提升观测可解释性。

原文摘要 · Abstract (English)

Generative AI models offer powerful capabilities but often lack transparency, making it difficult to interpret their output. This is critical in cases involving artistic or copyrighted content. This work introduces a search-inspired approach to improve the interpretability of these models by analysing the influence of training data on their outputs. Our method provides observational interpretability by focusing on a model's output rather than on its internal state. We consider both raw data and latent-space embeddings when searching for the influence of data items in generated content. We evaluate our method by retraining models locally and by demonstrating the method's ability to uncover influential subsets in the training data. This work lays the groundwork for future extensions, including user-based evaluations with domain experts, which is expected to improve observational interpretability further.

可解释性生成模型数据影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。