测试大模型对隐形Unicode指令的响应,发现工具使用会显著提升漏洞风险。
Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction Injection
- 用不可见的Unicode字符编码指令,测试模型是否无意识执行。
- 开启工具后模型合规率最高提升95个百分点,效应量达1.37。
- 不同厂商模型偏好不同编码方式,存在明显安全差异。
我们提出Reverse CAPTCHA评估框架,测试大语言模型是否遵循嵌入在看似正常文本中的隐形Unicode指令。与传统CAPTCHA区分人机不同,该基准利用模型能感知人类无法察觉的Unicode控制字符这一能力差。在两家厂商的五款模型上,测试了零宽二进制和Unicode标签两种编码方案,四种提示强度,两种载荷格式,并开启或关闭工具使用。共分析8,308条模型输出,结果显示:启用工具使模型遵从性大幅提升(Cohen's h最高达1.37,属大效应);模型表现出厂商特异性编码偏好(OpenAI模型偏爱零宽二进制,Anthropic模型倾向Unicode标签);显式解码指令可使单个模型在单一编码下合规率提升最高95个百分点。所有模型间的差异均具统计显著性(p < 0.05,Bonferroni校正)。结果揭示了通过隐形Unicode载荷进行提示注入这一被忽视的攻击面。
原文摘要 · Abstract (English)
We introduce Reverse CAPTCHA, an evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in otherwise normal-looking text. Unlike traditional CAPTCHAs that distinguish humans from machines, our benchmark exploits a capability gap: models can perceive Unicode control characters that are invisible to human readers. We evaluate five models from two providers across two encoding schemes (zero-width binary and Unicode Tags), four hint levels, two payload framings, and with tool use enabled or disabled. Across 8,308 model outputs, we find that tool use dramatically amplifies compliance (Cohen's h up to 1.37, a large effect), that models exhibit provider-specific encoding preferences (OpenAI models decode zero-width binary; Anthropic models prefer Unicode Tags), and that explicit decoding instructions increase compliance by up to 95 percentage points within a single model and encoding. All pairwise model differences are statistically significant (p < 0.05, Bonferroni-corrected). These results highlight an underexplored attack surface for prompt injection via invisible Unicode payloads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。