AI安全评估存在盲区,自动化测试无法发现未预设的威胁。
A Translational Note on AI Safety Evaluation

- 用预设威胁集评估安全,导致遗漏非预设风险
- 实测发现非英语提示下新威胁仍可被触发
- 需来自不同场景的评估者才能覆盖真实风险
近期研究显示,自动化红队测试在标准AI安全基准上比人工测试发现更多漏洞且成本更低,一些人据此认为人工评估正变得不再必要。但该比较衡量的是攻击者对预设危害集的搜索深度,而开发者未包含的危害对任何攻击者(无论是否自动化)都是不可见的。这一盲区在学术密码学和临床药物试验中也曾出现,评估虽内部有效,却忽略了其未指向的群体。我们称此为‘威胁模型覆盖缺口’,并在当前开源大模型中验证了其存在:非英语提示下暴露的新危害未被英语基准捕捉。填补该缺口需要评估者所处部署环境与开发者不同。这一需求是方法论层面的,根植于覆盖性,现有评估框架难以自行产生此类评估者。
原文摘要 · Abstract (English)
Recent studies report that automated red-teaming finds more vulnerabilities, at lower cost, than human red-teaming on standard AI safety benchmarks, and some read this as evidence that human evaluators are becoming dispensable. The comparison measures one thing and the conclusion claims another. A benchmark measures how thoroughly an attacker searches a predefined set of harms, fixed in advance by the developers, and a harm left out of that set is invisible to any attacker working inside it, automated or not. The same blind spot appeared in academic cryptography and in clinical drug trials, where an evaluation that was internally valid stayed silent about the population it was never pointed at. We call the AI-safety version the \emph{threat-model coverage gap}, and find that it persists in a current open-weight model, where harms surface in non-English prompts that English benchmarks miss. Closing it requires evaluators whose deployment context differs from the developers'. The case for those evaluators is methodological, grounded in coverage, and the existing evaluation frame is unlikely to produce them on its own.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。