arXiv:2502.06658cs.LG2025-02

通过引导生成数据,揭示模型在特定行为下的输入分布。

Guided Data Generation for Understanding Model Behavior

  • 用引导函数定义问题,生成满足特定行为的输入数据
  • 可生成模型预测标签、模型间分歧、敏感性及高风险输入分布
  • 适用于各类分类回归任务和模型类型,适合模型调试与分析

我们提出一种在输入空间生成分布的方法,作为理解训练模型行为的分析工具。该框架通过提出如“哪些输入会使模型表现出特定行为?”的问题,并以引导函数编码每个问题。生成的数据能揭示模型的行为模式。为展示该框架,我们提出了生成模型预测指定标签、两模型产生分歧、输出对参数扰动敏感、以及预测存在风险的输入分布等查询。该方法适用于多种分类与回归任务,可应用于不同类型的模型。

原文摘要 · Abstract (English)

We propose a method for generating distributions over the input space as an inspection tool for understanding trained models. Our framework poses questions of the form ``which inputs would make a trained model exhibit a specified behavior?'' and encodes each question through a guidance function. The generated data provide insights into how the models behave. To showcase our framework, we pose queries such as generating distributions of data where a specified label would be predicted by the model, where two distinct models would disagree, where the output is sensitive to parameter perturbations, and where predictions would be risky. Our method can be applied with a variety of classification and regression tasks and on a range of model types.

模型解释数据生成行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。