用扩散模型生成高风险且类别一致的样本,提升模型鲁棒性测试能力。
Generating Risky Samples with Conformity Constraints via Diffusion Models
- 利用图文嵌入与显式一致性评分约束类别符合性
- 生成样本风险度显著高于现有方法,且类别一致性更强
- 适合用于模型安全评估与训练数据增强
尽管神经网络在诸多任务中表现优异,但在面对某些样本时仍可能失效,带来应用风险。以往方法通过搜索现有数据集中的风险模式或注入扰动来发现风险样本,但受限于数据集覆盖范围,生成样本多样性不足。近期研究采用扩散模型生成超出原有数据集覆盖范围的新风险样本,但面临生成样本与目标类别一致性差的问题,易引入标签噪声,限制实际应用效果。为此,我们提出RiskyDiff,将文本与图像嵌入作为隐式类别一致性约束,并设计一致性评分以显式强化类别符合性,同时引入嵌入筛选与风险梯度引导机制,提升生成样本的风险程度。大量实验表明,RiskyDiff在风险程度、生成质量及类别一致性方面均显著优于现有方法。此外,实证显示,使用高一致性生成样本增强训练数据,可有效提升模型泛化能力。
原文摘要 · Abstract (English)
Although neural networks achieve promising performance in many tasks, they may still fail when encountering some examples and bring about risks to applications. To discover risky samples, previous literature attempts to search for patterns of risky samples within existing datasets or inject perturbation into them. Yet in this way the diversity of risky samples is limited by the coverage of existing datasets. To overcome this limitation, recent works adopt diffusion models to produce new risky samples beyond the coverage of existing datasets. However, these methods struggle in the conformity between generated samples and expected categories, which could introduce label noise and severely limit their effectiveness in applications. To address this issue, we propose RiskyDiff that incorporates the embeddings of both texts and images as implicit constraints of category conformity. We also design a conformity score to further explicitly strengthen the category conformity, as well as introduce the mechanisms of embedding screening and risky gradient guidance to boost the risk of generated samples. Extensive experiments reveal that RiskyDiff greatly outperforms existing methods in terms of the degree of risk, generation quality, and conformity with conditioned categories. We also empirically show the generalization ability of the models can be enhanced by augmenting training data with generated samples of high conformity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。