发现对抗攻击下矛盾类样本更稳定,暗示模型可能依赖特定推理模式。
Unpacking the Resilience of SNLI Contradiction Examples to Attacks
- 用通用对抗攻击测试模型脆弱性,揭示其对不同类别响应差异。
- 矛盾类准确率下降幅度显著低于蕴含与中立类,体现异常稳健性。
- 引入对抗样本微调可恢复模型性能,提升鲁棒性,适合关注模型可靠性研究者。
预训练模型在SNLI和MultiNLI等自然语言推理基准上表现优异,但其真实语言理解能力仍存疑。仅用假设句和标签训练的模型也能达到高准确率,表明其依赖数据集偏差和虚假相关性。我们采用通用对抗攻击来检验模型的脆弱性,分析显示蕴含和中立类准确率大幅下降,而矛盾类下降幅度较小。对包含对抗样本的增强数据集进行微调后,模型在标准集和挑战集上的性能均恢复至接近基线水平。研究结果表明,对抗触发器有助于识别虚假相关性并提升模型鲁棒性,同时揭示了矛盾类对对抗攻击的异常韧性。
原文摘要 · Abstract (English)
Pre-trained models excel on NLI benchmarks like SNLI and MultiNLI, but their true language understanding remains uncertain. Models trained only on hypotheses and labels achieve high accuracy, indicating reliance on dataset biases and spurious correlations. To explore this issue, we applied the Universal Adversarial Attack to examine the model's vulnerabilities. Our analysis revealed substantial drops in accuracy for the entailment and neutral classes, whereas the contradiction class exhibited a smaller decline. Fine-tuning the model on an augmented dataset with adversarial examples restored its performance to near-baseline levels for both the standard and challenge sets. Our findings highlight the value of adversarial triggers in identifying spurious correlations and improving robustness while providing insights into the resilience of the contradiction class to adversarial attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。