arXiv:2510.24446cs.CLcs.CV2025-10Conference of the …被引 1

用黑盒对抗改写测试文本模型在语义不变下的分割鲁棒性

SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space

  • 在文本自编码器潜空间中通过强化学习生成语法正确且语义一致的对抗改写
  • 在ReasonSeg和LLMSeg-40k数据集上成功率提升至原有方法的2倍
  • 揭示先进模型对语义等价改写仍脆弱,适合关注模型安全与鲁棒性的研究者

多模态大语言模型在视觉-语言任务中表现出色,例如基于文本查询生成分割掩码的推理分割。现有工作主要关注图像输入扰动,而真实场景中用户以不同方式表达相同意图的语义等价文本改写仍未被充分探索。为此,我们提出一种新型对抗改写任务:生成语法正确且保留原意的改写,同时降低分割性能。为评估对抗改写质量,我们构建了经人工验证的自动评估协议。此外,我们提出SPARTA——一种基于文本自编码器低维语义潜空间、黑盒句级优化的方法,由强化学习驱动。SPARTA在ReasonSeg和LLMSeg-40k数据集上的成功率比之前方法最高提升2倍。我们使用SPARTA与基线方法评估先进推理分割模型的鲁棒性,发现其在严格语义和语法约束下仍易受对抗改写影响。所有代码与数据将在论文接收后公开。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have shown impressive capabilities in vision-language tasks such as reasoning segmentation, where models generate segmentation masks based on textual queries. While prior work has primarily focused on perturbing image inputs, semantically equivalent textual paraphrases-crucial in real-world applications where users express the same intent in varied ways-remain underexplored. To address this gap, we introduce a novel adversarial paraphrasing task: generating grammatically correct paraphrases that preserve the original query meaning while degrading segmentation performance. To evaluate the quality of adversarial paraphrases, we develop a comprehensive automatic evaluation protocol validated with human studies. Furthermore, we introduce SPARTA-a black-box, sentence-level optimization method that operates in the low-dimensional semantic latent space of a text autoencoder, guided by reinforcement learning. SPARTA achieves significantly higher success rates, outperforming prior methods by up to 2x on both the ReasonSeg and LLMSeg-40k datasets. We use SPARTA and competitive baselines to assess the robustness of advanced reasoning segmentation models. We reveal that they remain vulnerable to adversarial paraphrasing-even under strict semantic and grammatical constraints. All code and data will be released publicly upon acceptance.

推理分割对抗攻击文本改写鲁棒性评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。