警告:别再称‘正向后门’,秘密对齐需严格验证
Position: Retire the "Positive Backdoor" Label -- Secret Alignment Requires Strict and Systematic Evaluation

- 将隐藏行为统一为‘秘密对齐’,强调安全需实证
- 六项评估显示触发机制在保密性、完整性上极脆弱
- 适合关注模型安全与合规的开发者和研究者
本文主张学术界应摒弃‘正向后门’标签,改用‘秘密对齐’描述触发激活的隐秘行为。在开放权重大模型时代,语言模型成为私有数字资产,面临未经授权访问、模型窃取和行为滥用等风险。近期多项工作以‘正向后门’名义提出解决方案,我们将其统一为三类应用:访问控制、所有权归属与安全约束。通过在六项核心属性(有效性、无害性、持久性、效率、鲁棒性、可靠性)上评估三个代表性方案,发现触发-行为映射普遍存在脆弱性,尤其在保密性、完整性和可用性(CIA)方面表现不佳。我们进一步关联行为密度与决策复杂度,提供理解部署风险的新视角,呼吁建立全社区标准化评估体系,使‘秘密对齐’声明可被验证。
原文摘要 · Abstract (English)
This position paper argues that the AI/ML community should stop overclaiming and retire the label "positive backdoor," and instead treat trigger-activated hidden behaviors as Secret Alignment. Crucially, protective claims based on Secret Alignment should be presumed not secure by default unless supported by rigorous, standardized evaluation. The Private AI era, enabled by open-weight LLMs and accessible training/inference stacks, turns language models into privately owned digital assets, creating security concerns around unauthorized access, model theft, and behavioral misuse. Recently, a line of work framed as "positive backdoors" has been proposed to address these challenges. To ground our position in evidence, we unify these proposals as covert trigger-behavior associations for access gating, ownership attribution, and safety enforcement, and evaluate three representative applications across six core properties: effectiveness, harmlessness, persistence, efficiency, robustness, and reliability. Our results reveal substantial brittleness - especially in the confidentiality, integrity, and availability (CIA) - of trigger-behavior mappings often underrepresented by existing claims. We further relate these outcomes to behavior density and decision complexity, offering a behavioral lens for understanding deployment-time risks and motivating community-wide evaluation that makes Secret Alignment claims provable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。