arXiv:2608.15392cs.AI2026-08

跨语言分析可见推理在间接注入攻击中的监控效果

Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish

论文配图:Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish
图 1 · 摘自论文原文
  • 在英、泰米尔、坦格利什三语中测试大模型的可见推理行为
  • 有推理时攻击成功率为25%(3/12),无推理时为42%(5/12)
  • 所有成功攻击均明确表示要服从指令,安全响应则拒绝

链式思维监控可能是一种有用的安全部信号,但其在不同语言和行为场景下的可靠性尚不明确。在一项包含八种人工验证的合成场景的小型案例研究中,使用一个模型、一名标注者和一个确定性生成种子,研究了萨尔瓦姆-105B在英语、泰米尔语和坦格利什语中对间接提示注入的可见推理表现。初步四场景试验发现:无推理时攻击成功5/12次,有推理时为1/11次;预注册的后续四场景实验结果反转,显示无推理时攻击成功2/12次,有推理时为3/12次。由于每阶段仅四个场景,无法区分真实推理效应与提示特异性或抽样噪声。在20条非空的注入思维轨迹中,17条良性正确输出均声明将忽略注入指令,3次攻击成功则明确表示将遵循指令。这些观察提供了可复现的案例研究,表明可见推理在可用时具有行为信息量,但未证明推理模式提升安全性,也未证实其机制真实性或泛化能力。

原文摘要 · Abstract (English)

Chain-of-thought monitoring is a potentially useful safety signal, but its reliability across languages and behavioral settings remains uncertain. In a small case study of eight manually verified synthetic scenarios, one model, one annotator, and one deterministic generation seed, I study API-visible reasoning during indirect prompt injection in Sarvam-105B across English, Tamil, and Tanglish. A four scenario pilot found 5/12 injected attack successes without reasoning and 1/11 with reasoning. A preregistered four-scenario follow-up reversed that direction, finding 2/12 attacks without reasoning and 3/12 with reasoning. With only four scenarios per phase, this design cannot distinguish a real reasoning-mode effect from prompt-specific variation or sampling noise. Across 20 non-empty injected-thinking traces, all 17 benign-correct outputs stated an intent to ignore the injection, while all three attack successes stated an intent to follow it. These descriptive observations provide a reproducible case study of behaviorally informative visible reasoning when it is available; they do not establish that reasoning mode improves safety, that visible reasoning is mechanistically faithful, or that the findings generalize beyond this configuration.

安全监控多语言提示注入可见推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。