arXiv:2510.13570cs.LG2025-10

提出可精准干扰特定大模型的对抗攻击,动摇排行榜公正性

Selective Adversarial Attacks on LLM Benchmarks

  • 设计选择性攻击协议,用约束和替代模型生成针对性扰动
  • 在MMLU基准上实现对特定模型性能的显著降低,不影响其他模型
  • 揭示排行榜评估存在隐蔽偏差,适合关注模型评测公平性的研究者

基准测试结果日益决定大模型的信任度、选型与部署,但这些评估仍易受语义等价的对抗扰动影响。现有NLP对抗鲁棒性研究多聚焦于普遍性文本攻击,未解决能否仅针对特定模型降级或提升性能的问题。本文形式化该问题,在广泛使用的MMLU基准(衡量模型跨学科知识与推理能力)上开展研究。通过集成TextAttack框架中的经典攻击方法,构建选择性评估协议,设计定制化约束以增强攻击选择性,并提出基于替代大模型的生成管道。实验证明,选择性对抗攻击确实存在,可显著改变模型相对排名,挑战排行榜驱动评估的公平性、可复现性与透明性。结果呼吁在评测中引入扰动感知报告与鲁棒性诊断,表明微小编辑即可扭转比较结论。

原文摘要 · Abstract (English)

Benchmarking outcomes increasingly govern trust, selection, and deployment of LLMs, yet these evaluations remain vulnerable to semantically equivalent adversarial perturbations. Prior work on adversarial robustness in NLP has emphasized text attacks that affect many models equally, leaving open the question of whether it is possible to selectively degrade or enhance performance while minimally affecting other models. We formalize this problem and study selective adversarial attacks on MMLU - a widely used benchmark designed to measure a language model's broad general knowledge and reasoning ability across different subjects. Using canonical attacks integrated into TextAttack framework, we introduce a protocol for selectivity assessment, develop a custom constraint to increase selectivity of attacks and propose a surrogate-LLM pipeline that generates selective perturbations. Empirically, we find that selective adversarial attacks exist and can materially alter relative rankings, challenging the fairness, reproducibility, and transparency of leaderboard-driven evaluation. Our results motivate perturbation-aware reporting and robustness diagnostics for LLM evaluation and demonstrate that even subtle edits can shift comparative judgments.

对抗攻击大模型评测选择性攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。