用心理画像技术识别并清除模型中的隐蔽后门,保障AI安全
Hypnopaedia-Aware Machine Unlearning via Psychometrics of Artificial Mental Imagery
- 通过反向工程与统计推断检测隐藏的恶意触发模式
- 利用模型反演生成人工心理意象,打断错误优化路径
- 可主动识别后门感染概率,适合安全敏感场景使用
神经后门是潜在的网络安全漏洞,使学习系统易受未经授权的操控,可能导致人工智能被武器化。后门攻击在学习过程中秘密植入触发信号,如同催眠状态下向潜意识灌输思想。当特定感官刺激激活时,会引发条件反射,使机器执行预设行为。本研究提出一种动态监控后门威胁的控制论框架,针对不可信数据源设计自知型遗忘机制,可自主剥离模型对后门触发器的依赖。通过逆向工程与统计推断,检测欺骗性模式并估计后门感染概率。采用模型反演激发人工心理意象,利用随机过程干扰优化路径,避免收敛至潜在错误模式。随后进行假设分析,评估每个可疑模式作为真实触发器的可能性,并推断感染概率。研究核心目标是在知识保真度与后门脆弱性之间维持稳定平衡。
原文摘要 · Abstract (English)
Neural backdoors represent insidious cybersecurity loopholes that render learning machinery vulnerable to unauthorised manipulations, potentially enabling the weaponisation of artificial intelligence with catastrophic consequences. A backdoor attack involves the clandestine infiltration of a trigger during the learning process, metaphorically analogous to hypnopaedia, where ideas are implanted into a subject's subconscious mind under the state of hypnosis or unconsciousness. When activated by a sensory stimulus, the trigger evokes a conditioned reflex that directs a machine to mount a predetermined response. In this study, we propose a cybernetic framework for constant surveillance of backdoor threats, driven by the dynamic nature of untrustworthy data sources. We develop a self-aware unlearning mechanism to autonomously detach a machine's behaviour from the backdoor trigger. Through reverse engineering and statistical inference, we detect deceptive patterns and estimate the likelihood of backdoor infection. We employ model inversion to elicit artificial mental imagery, using stochastic processes to disrupt optimisation pathways and avoid convergent but potentially flawed patterns. This is followed by hypothesis analysis, which estimates the likelihood of each potentially malicious pattern as the true trigger and infers the probability of infection. The primary objective of this study is to maintain a stable state of equilibrium between knowledge fidelity and backdoor vulnerability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。