arXiv:2607.01518cs.CRcs.RO2026-07

攻击者用特定文字触发机器人视觉大模型过思考,导致决策延迟超6倍。

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems

论文配图:Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems
图 1 · 摘自论文原文
  • 通过文本片段挖掘引发过思考的关键词特征
  • 单个触发器使推理延迟最高达6.96倍,物理打印仍可致4.74倍延迟
  • 攻击对多模型有效,适合研究安全漏洞的开发者参考

大型视觉语言模型(LVLM)被广泛集成于机器人系统中。然而,这些模型可能表现出过思考行为,即生成过长的推理过程,导致严重推理延迟。攻击者可刻意嵌入人类可读的场景文本,触发目标机器人的过思考,造成决策延迟,构成严重安全隐患(即过思考诱导的延迟攻击)。本工作提出三阶段框架,系统识别并验证了可转移的过思考触发器:首先构建多样化推理密集型场景文本语料库,从短响应前缀中提取与过思考相关的词汇特征;其次在严格黑盒设置下,基于前缀代理得分高效搜索,并仅对少数候选进行完整延迟验证;最后在未见图像和多个LVLM上评估黑盒迁移性,报告延迟放大率与攻击成功率。在三个代表性LVLM上,所有触发器均导致延迟放大率超过1.0x,最强单一触发器达6.96x;物理打印文本仍可引发最高4.74x延迟放大。结果表明,所发现触发器可在多模型间迁移,持续造成机器人系统显著延迟。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have been increasingly integrated into robotic systems. However, these models may exhibit overthinking behaviors, where they generate excessively long reasoning traces, incurring an excessive inference time. This overthinking behavior poses a serious risk to robotic systems, as the adversary can deliberately trigger overthinking to slow down the decision making of a victim robotic system, causing a variety of safety issues (i.e., an overthinking-induced slowdown attack). To initiate this attack, an adversary can embed carefully crafted, human-readable scene text into the visual scene observed by a victim robotic agent, causing significant inference delays even under a strict black-box setting. Therefore, the embedded scene text serves as a significant "trigger" for the attack. This work systematically identifies and validates transferable triggers of overthinking in robotic systems by introducing a three-stage framework. First, we construct a diverse corpus of reasoning-intensive scene text and extract overthinking-correlated lexical features from short response prefixes. Second, we perform an efficient black-box search guided by a prefix-based proxy score while selectively confirming a small set of top candidates with full latency measurements. Third, we evaluate black-box transfer using a fixed pool of triggers on unseen images and multiple LVLMs, reporting latency amplification and attack success rates under standard thresholds. Across three representative LVLMs, all triggers yield slowdown ratios greater than 1.0x, with the strongest single-trigger case reaching 6.96x. The physical printing of the text trigger still causes up to 4.74x latency amplification. These results demonstrate that our discovered triggers are transferred between multiple LVLM models and consistently cause significant slowdowns in robotic systems.

视觉语言模型安全攻击机器人系统延迟攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。