arXiv:2605.13338cs.CRcs.AI2026-05中稿 · ICML被引 1

通过智能扰动输入诱导大模型过度思考,实现黑箱拒绝服务攻击

Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models

论文配图:Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models
图 1 · 摘自论文原文
  • 用分层遗传算法优化问题结构,自动诱发模型冗长推理
  • 在MATH数据集上使输出长度最长提升26.1倍,显著超过人工构造基线
  • 攻击样本对商用大模型具强迁移性,揭示推理模型共性漏洞

大型推理模型(LRMs)在需要多步推断的系统中日益普及,但其对计算资源的依赖暴露了新漏洞。当面对不完整或逻辑矛盾的输入时,LRMs易产生过度推理,导致响应延迟和能耗飙升,形成拒绝服务(DoS)式资源耗尽风险。本文研究此攻击面,提出一种黑箱自动化框架,通过系统性扰动输入逻辑结构诱导过度思考。方法采用分层遗传算法(HGA),在结构化问题分解上优化复合适应度函数,以最大化输出长度与反思型过思标记。在四个前沿推理模型上,该方法显著延长输出,于MATH基准上最高达26.1倍增幅,并持续优于良性及人工构造的缺失前提基线。进一步验证了攻击样本的强迁移性:由小型代理模型生成的对抗输入,仍对大型商业推理模型有效。结果表明,过度思考是现代推理系统的共性可被利用漏洞,亟需更鲁棒的防御机制。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) are increasingly integrated into systems requiring reliable multi-step inference, yet this growing dependence exposes new vulnerabilities related to computational availability. In particular, LRMs exhibit a tendency to "overthink", producing excessively long and redundant reasoning traces, when confronted with incomplete or logically inconsistent inputs. This behavior significantly increases inference latency and energy consumption, forming a potential vector for denial-of-service (DoS) style resource exhaustion. In this work, we investigate this attack surface and propose an automated black-box framework that induces overthinking in LRMs by systematically perturbing the logical structure of input problems. Our method employs a hierarchical genetic algorithm (HGA) operating on structured problem decompositions, and optimizes a composite fitness function designed to maximize both response length and reflective overthinking markers. Across four state-of-the-art reasoning models, the proposed method substantially amplifies output length, achieving up to a 26.1x increase on the MATH benchmark and consistently outperforming benign and manually crafted missing-premise baselines. We further demonstrate strong transferability, showing that adversarial inputs evolved using a small proxy model retain high effectiveness against large commercial LRMs. These findings highlight overthinking as a shared and exploitable vulnerability in modern reasoning systems, underscoring the need for more robust defenses.

大模型安全拒绝服务过度推理遗传算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。