arXiv:2605.25889cs.CRcs.LG2026-05

视觉语言动作模型的能力与鲁棒性存在理论极限,不可兼得。

Capability and Robustness Cannot Both Be Free: An Information-Theoretic Bound for Vision-Language-Action Models

  • 从信息论角度证明能力与鲁棒性之和受任务熵和对抗信道容量限制
  • 实测显示微小扰动使成功率从95%暴跌至5%以下,验证理论边界
  • 提供无需标签的诊断工具,适用于多种模型架构和任务类型

视觉语言动作(VLA)模型在干净输入下表现优异,但在微小对抗扰动下迅速失效:$16/255$ 的 PGD 攻击使 OpenVLA-7B 在 LIBERO 任务上的成功率从 95% 降至不足 5%。这种能力与鲁棒性的权衡是否存在理论下限曾是未解之谜。本文证明其存在:任意 VLA 策略的能力 $I( ext{A}^*; ext{A}_ ext{pi})$ 与鲁棒性 $I( ext{A}_ ext{pi}; ilde{ ext{A}}_ ext{pi}) - I( ext{A}_ ext{pi};oldsymbol{ u})$ 之和不超过 $H( ext{A}^*) + I(X; ilde{X})$,即任务熵加上对抗信道容量。证明基于两次数据处理不等式。像素级边界松散约 $10^3$ 纳特,作为上限保证;编码器特定推论将其收紧一个数量级以上,现实能力已消耗 $5$–$9\/%$ 预算。在 $308$ 个测试单元中零违规:包括 $252$ 个闭式高斯-VLA、$48$ 个 OpenVLA-7B+LIBERO+PGD($4$ 套件 × $4$ ε × $3$ 种种子)、$4$ 个 Square-Attack 和 $4$ 个多步攻击($T=10$)。互补可测量不等式 $\Rob_{\text{disc}} \le \Cap_{\text{disc}}$ 在 $144$ 个跨架构单元中成立,涵盖 OpenVLA、OpenVLA-OFT(连续-$L_1$)和 SmolVLA(流匹配)。同一构造衍生出三项无标签诊断:预飞行编码器上限、防御溯源探针(定位输入侧或语言模型干预)、头无关鲁棒性比率,可在离散词元、$L_1$ 回归和流匹配策略间直接比较。这些共同填补了当前跨设置防御与架构比较的空白。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models reach high success rates on clean inputs but collapse under small adversarial perturbations: a $16/255$ PGD attack drops OpenVLA-7B's LIBERO success from $95\%$ to under $5\%$. Whether this trade-off has a theoretical floor was open. We prove that it does. For any VLA policy, capability $I(\Astar;\Api)$ and robustness $I(\Api;\Atildepi)-I(\Api;δ)$ sum to at most $H(\Astar)+I(X;\Xtilde)$, the task entropy plus adversarial channel capacity. The proof reduces to two applications of the Data Processing Inequality. The pixel-level bound is loose by $\sim 10^3$ nats and serves as a ceiling guarantee; an encoder-specific corollary tightens it by over an order of magnitude, into a regime where realized capability already consumes $5$--$9\%$ of the budget. We validate Theorem~\ref{thm:main} with zero violations across $308$ cells: $252$ closed-form Gaussian-VLA, $48$ OpenVLA-7B$+$LIBERO$+$PGD ($4$ suites $\times$ $4$ $\eps$ $\times$ $3$ seeds), $4$ Square-Attack, and $4$ multi-step ($T{=}10$). A complementary measurability inequality $\Rob_{\text{disc}} \le \Cap_{\text{disc}}$ further holds across $144$ cross-architecture cells spanning OpenVLA, OpenVLA-OFT (continuous-$L_1$), and SmolVLA (flow-matching). The same construction yields three label-free diagnostics: a pre-flight encoder ceiling, a defense-forensics probe that localizes input-side vs.\ language-model intervention, and a head-agnostic robustness ratio comparable across discrete-token, $L_1$-regression, and flow-matching policies. Together these provide the cross-setting axis defense and architecture comparisons currently lack.

视觉语言动作对抗鲁棒性信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。