arXiv:2504.21042cs.CRcs.AI2025-04中稿 · The ACM Conference…

通过概念漂移分析,揭示AI训练与推理中的安全漏洞与偏见来源。

What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift

  • 利用预训练多模态模型检测概念漂移,定位数据污染根源。
  • 发现隐蔽广告、隐私泄露及训练数据不平衡等多重风险。
  • 适合关注AI安全、可解释性与公平性的研究人员和开发者。

人工智能的广泛应用加剧了对可信度的担忧,包括完整性、隐私性、鲁棒性和偏差问题。为此,我们提出ConceptLens框架,利用预训练多模态模型分析探测样本中的概念漂移,以识别完整性威胁的根本原因。该框架在检测原始数据投毒攻击方面表现优异,并揭示了隐含偏差注入的风险,例如通过恶意概念漂移生成隐蔽广告。它能在不修改数据的情况下识别高风险样本并提前过滤,揭示由训练数据不完整或不平衡引发的模型弱点。在模型层面,能追踪目标模型过度依赖的概念,识别误导性概念,并解释破坏关键概念如何损害模型性能。此外,还发现了生成内容中的社会学偏见,揭示不同社会背景下的差异。令人震惊的是,即使训练和推理数据看似安全,仍可能被无意中利用,从而削弱安全对齐效果。本研究为提升AI系统可信度提供了可操作的洞见,有助于加速技术采纳与创新。

原文摘要 · Abstract (English)

The growing adoption of artificial intelligence (AI) has amplified concerns about trustworthiness, including integrity, privacy, robustness, and bias. To assess and attribute these threats, we propose ConceptLens, a generic framework that leverages pre-trained multimodal models to identify the root causes of integrity threats by analyzing Concept Shift in probing samples. ConceptLens demonstrates strong detection performance for vanilla data poisoning attacks and uncovers vulnerabilities to bias injection, such as the generation of covert advertisements through malicious concept shifts. It identifies privacy risks in unaltered but high-risk samples, filters them before training, and provides insights into model weaknesses arising from incomplete or imbalanced training data. Additionally, at the model level, it attributes concepts that the target model is overly dependent on, identifies misleading concepts, and explains how disrupting key concepts negatively impacts the model. Furthermore, it uncovers sociological biases in generative content, revealing disparities across sociological contexts. Strikingly, ConceptLens reveals how safe training and inference data can be unintentionally and easily exploited, potentially undermining safety alignment. Our study informs actionable insights to breed trust in AI systems, thereby speeding adoption and driving greater innovation.

AI安全概念漂移可解释性偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。