为动态AI系统建立可追溯的安全评估新范式
Referential Security as a New Paradigm for AI Evaluations
- 以可验证的系统标识取代静态标签,确保评估对象明确
- 实现评估结果可复现、审计可追踪、跨平台可比
- 适合监管机构、安全审计员及持续迭代模型团队
安全评估依赖稳定的标识符。当前人工智能系统持续更新,虽模型名称不变,但权重、提示词、检索机制、滥用检测器、推理设置和部署架构常在未通知情况下变更,导致现有评估多针对表面标签而非具体可识别系统。为此,我们提出“指称安全性”新范式:安全问题不仅在于模型是否安全,更在于后续各方能否确凿判定某项安全声明所针对的具体系统。该范式将模型身份重构为可实证属性,使指称稳定性与安全主张分离。此框架提升了三类关键流程的可行性:可复现评估、长期审计有效性和跨厂商等价性。通过基于可验证实体的评估,确保安全审计与监管结论在动态系统的全生命周期中保持实际效用。
原文摘要 · Abstract (English)
Security evaluations inherently depend on stable identifiers. Any finding, audit, or regulatory decision must remain attached to the specific artifact it pertains to. Continuously updated artificial intelligence systems violate this core assumption, with public model designations remaining static while underlying weights, prompts, retrieval mechanisms, misuse classifiers, inference settings, and serving infrastructures undergo unannounced modifications. Consequently, current evaluations frequently apply to superficial labels rather than identifiable and distinct systems. To resolve this, we propose referential security as a new paradigm for AI evaluation. The fundamental security question extends beyond whether a model is safe to whether subsequent parties can conclusively determine which system a specific safety claim addressed. This approach reframes model identity as an empirically verifiable property and separates referential stability from the substantive security claims it conditions. This framework brings tractability to three critical workflows that current practices handle poorly. Specifically, it enables reproducible evaluation, longitudinal audit validity, and cross-provider equivalence. By grounding these evaluations in verifiable artifacts, our approach ensures that safety audits and regulatory findings maintain their empirical utility across the operational lifecycle of dynamic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。