arXiv:2605.18593cs.CRcs.AI2026-05

物理文字干扰可让家用机器人拿错物品,威胁安全

Not What You Asked For: Typographic Attacks in Household Robot Manipulation

  • 用贴纸诱导视觉语言模型误判物体,保持几何定位不变
  • 攻击成功率高达67.8%,在成功任务中达70.0%
  • 错误识别会引发真实抓取动作,影响实际操作安全

开放词汇具身智能体越来越多依赖CLIP等视觉-语言模型进行物体感知与任务理解。然而,共享嵌入空间带来的灵活性也引入了结构化漏洞:物理场景中的印刷文字可语义覆盖视觉判断。尽管已有研究在静态2D基准和3D导航任务中量化了此类威胁,但其对家用机器人完整感知-规划-执行流程的影响仍未被探索。本文基于Habitat仿真环境与HomeRobot基准评估了该攻击。我们提出解耦感知架构,将冻结的CLIP编码器暴露于对抗性贴纸,同时通过DETIc维持几何定位。在59个可归因的测试场景中,攻击整体成功率达67.8%,在完全成功任务中上升至70.0%,且在无视角控制、遮挡及无感知优化条件下仍有效。关键发现是:感知错误会通过持久的3D语义地图传播,导致动力学失败——即由被污染的语义状态驱动,机器人实际抓取并运送错误物体。这表明文字攻击不仅是分类错误,更是可引发真实物理失误的安全威胁。

原文摘要 · Abstract (English)

Open-vocabulary embodied AI agents increasingly rely on vision-language models such as CLIP for object perception and task grounding. However, the shared embedding space that enables this flexibility introduces a structural vulnerability to typographic attacks, where printed text in a physical scene semantically overrides visual judgment. While prior work has quantified this threat in static 2D benchmarks and 3D navigation tasks, its impact on the full Sense-Plan-Act pipeline of household robot manipulation remains unexplored. This work evaluates typographic attacks in a Habitat-based simulation using the HomeRobot benchmark. We introduce a decoupled perception architecture that exposes a frozen CLIP encoder to adversarial stickers while maintaining geometric grounding via DETIC. In a controlled evaluation pool of 59 attributable episodes, the attack achieves an overall Attack Success Rate (ASR) of 67.8%, rising to 70.0% among fully successful episodes, under uncontrolled viewing angles and occlusion with no perceptual optimization. Critically, we find that perceptual errors propagate through the persistent 3D semantic map to produce kinetic failures, defined here as physically executed grasping and transport of the wrong object driven by an adversarially poisoned semantic state. In these cases, the robot physically grasps and delivers the wrong object to a target receptacle. These results establish typographic misclassification as a real, measurable, and physically consequential threat to the safety of modular manipulation pipelines that prior typographic attack research has left unexamined.

机器人安全对抗攻击视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。