arXiv:2507.08983cs.LGcs.CR2025-07被引 5

攻击者可利用排行榜隐蔽分发带毒模型,伪装成高性能模型。

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models

  • 通过排行榜机制注入恶意功能,同时保持高排名表现。
  • 在文本、语音、图像等4种模态中成功实现高排名与恶意行为共存。
  • 提醒研究者警惕未验证来源模型,需重构评估机制防范风险。

尽管对机器学习模型的投毒攻击已有广泛研究,但攻击者如何大规模分发被污染模型的机制仍鲜有探索。本文揭示了模型排行榜——作为模型发现与评估的排名平台——可能成为攻击者隐蔽分发有毒模型的强大渠道。我们提出TrojanClimb框架,可在维持领先排行榜性能的同时注入恶意行为。实验覆盖文本嵌入、文本生成、文本转语音和文本转图像四种模态,证明攻击者可实现高排名并嵌入任意有害功能,如后门或偏见注入。研究揭示了机器学习生态中的重大漏洞,亟需重构排行榜评估机制以检测并过滤恶意(如中毒)模型,同时也暴露了从非可信来源采纳模型带来的广泛安全风险。

原文摘要 · Abstract (English)

While poisoning attacks on machine learning models have been extensively studied, the mechanisms by which adversaries can distribute poisoned models at scale remain largely unexplored. In this paper, we shed light on how model leaderboards -- ranked platforms for model discovery and evaluation -- can serve as a powerful channel for adversaries for stealthy large-scale distribution of poisoned models. We present TrojanClimb, a general framework that enables injection of malicious behaviors while maintaining competitive leaderboard performance. We demonstrate its effectiveness across four diverse modalities: text-embedding, text-generation, text-to-speech and text-to-image, showing that adversaries can successfully achieve high leaderboard rankings while embedding arbitrary harmful functionalities, from backdoors to bias injection. Our findings reveal a significant vulnerability in the machine learning ecosystem, highlighting the urgent need to redesign leaderboard evaluation mechanisms to detect and filter malicious (e.g., poisoned) models, while exposing broader security implications for the machine learning community regarding the risks of adopting models from unverified sources.

模型安全投毒攻击排行榜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。