提出针对语言视觉关联追踪系统的新型攻击框架,揭示其安全漏洞。
See No Evil: Adversarial Attacks Against Linguistic-Visual Association in Referring Multi-Object Tracking Systems
- 设计对抗性扰动破坏语言-视觉关联与目标匹配逻辑
- 物理和数字扰动均能引发跟踪错误和ID切换
- 适用于安全评估与鲁棒性设计的先进追踪系统
语言-视觉理解推动了先进感知系统的发展,尤其是新兴的指代多目标追踪(RMOT)范式。通过自然语言查询,RMOT系统可选择性追踪满足语义描述的对象,由基于Transformer的时空推理模块引导。端到端(E2E)RMOT模型将特征提取、时序记忆与空间推理统一于Transformer主干中,实现对融合文本-视觉表示的长程时空建模。尽管如此,RMOT系统的可靠性与鲁棒性仍缺乏深入研究。本文从设计逻辑角度审视RMOT系统的安全性,识别出影响语言-视觉指代与目标匹配组件的对抗性漏洞。此外,我们发现采用基于FIFO内存的先进模型存在新漏洞:对时空推理的定向且一致攻击会导致错误在历史缓冲区中持续存在多个后续帧。我们提出VEIL框架,旨在破坏RMOT模型的统一指代-匹配机制。实验表明,精心设计的数字与物理扰动可破坏追踪逻辑可靠性,引发跟踪ID切换与终止。我们在Refer-KITTI数据集上进行综合评估,验证了VEIL的有效性,凸显了在大规模关键应用中亟需安全感知的RMOT设计。
原文摘要 · Abstract (English)
Language-vision understanding has driven the development of advanced perception systems, most notably the emerging paradigm of Referring Multi-Object Tracking (RMOT). By leveraging natural-language queries, RMOT systems can selectively track objects that satisfy a given semantic description, guided through Transformer-based spatial-temporal reasoning modules. End-to-End (E2E) RMOT models further unify feature extraction, temporal memory, and spatial reasoning within a Transformer backbone, enabling long-range spatial-temporal modeling over fused textual-visual representations. Despite these advances, the reliability and robustness of RMOT remain underexplored. In this paper, we examine the security implications of RMOT systems from a design-logic perspective, identifying adversarial vulnerabilities that compromise both the linguistic-visual referring and track-object matching components. Additionally, we uncover a novel vulnerability in advanced RMOT models employing FIFO-based memory, whereby targeted and consistent attacks on their spatial-temporal reasoning introduce errors that persist within the history buffer over multiple subsequent frames. We present VEIL, a novel adversarial framework designed to disrupt the unified referring-matching mechanisms of RMOT models. We show that carefully crafted digital and physical perturbations can corrupt the tracking logic reliability, inducing track ID switches and terminations. We conduct comprehensive evaluations using the Refer-KITTI dataset to validate the effectiveness of VEIL and demonstrate the urgent need for security-aware RMOT designs for critical large-scale applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。