通过生成对抗性3D物体,破坏视觉语言导航系统的可靠性。
Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion
- 构建对抗性3D物体,融合物理属性与视觉语言模型感知进行协同优化。
- 在多视角下迭代融合,使导航代理性能显著下降,但对正常任务干扰小。
- 揭示了视觉语言模型导航服务的安全隐患,适合安全与系统鲁棒性研究者。
我们提出对抗性物体融合(AdvOF),一种针对服务环境中基于视觉-语言导航(VLN)代理的新攻击框架,通过生成对抗性3D物体实现。尽管大型语言模型(LLMs)和视觉语言模型(VLMs)提升了服务导向导航系统的感知与决策能力,但其集成带来了关键任务流程中的安全隐患。现有攻击方法未考虑服务计算场景中对可靠性和服务质量(QoS)的高要求。AdvOF首先在2D与3D空间中精确聚合并对齐目标物体位置,定义并渲染对抗性物体;随后,通过在物理属性与VLM感知间施加正则化,协同优化对抗性物体。通过为不同视角分配重要性权重,利用局部更新与验证进行多视角、稳定的迭代融合。大量实验表明,AdvOF可在对抗条件下有效降低代理性能,同时对正常导航任务干扰极小。本工作深化了对基于VLM导航系统服务安全性的理解,为物理世界部署中的鲁棒服务组合提供了计算基础。
原文摘要 · Abstract (English)
We present Adversarial Object Fusion (AdvOF), a novel attack framework targeting vision-and-language navigation (VLN) agents in service-oriented environments by generating adversarial 3D objects. While foundational models like Large Language Models (LLMs) and Vision Language Models (VLMs) have enhanced service-oriented navigation systems through improved perception and decision-making, their integration introduces vulnerabilities in mission-critical service workflows. Existing adversarial attacks fail to address service computing contexts, where reliability and quality-of-service (QoS) are paramount. We utilize AdvOF to investigate and explore the impact of adversarial environments on the VLM-based perception module of VLN agents. In particular, AdvOF first precisely aggregates and aligns the victim object positions in both 2D and 3D space, defining and rendering adversarial objects. Then, we collaboratively optimize the adversarial object with regularization between the adversarial and victim object across physical properties and VLM perceptions. Through assigning importance weights to varying views, the optimization is processed stably and multi-viewedly by iterative fusions from local updates and justifications. Our extensive evaluations demonstrate AdvOF can effectively degrade agent performance under adversarial conditions while maintaining minimal interference with normal navigation tasks. This work advances the understanding of service security in VLM-powered navigation systems, providing computational foundations for robust service composition in physical-world deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。