用树搜索强化视觉语言模型安全,防跨模态攻击。
VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search
- 通过视觉文本交互提示嵌入安全约束,引导推理过程。
- 采用MCTS生成多样化安全敏感提示路径,提升风险识别率。
- 支持实时风险检测,适合高安全需求的多模态应用。
大型视觉语言模型(LVLMs)在多模态感知与生成方面取得显著进展,但其安全对齐仍面临严峻挑战。现有方法易受多模态越狱攻击,因视觉输入引入新攻击面,推理链缺乏安全监督,且对齐在模态融合下常退化。为此,我们提出VisuoAlign,一种基于提示引导的多模态树搜索安全对齐框架。该框架通过视觉-文本交互提示将安全约束嵌入推理过程,采用蒙特卡洛树搜索(MCTS)系统构建多样化的安全敏感提示轨迹,并引入基于提示的缩放机制,实现实时风险检测与合规响应。大量实验表明,VisuoAlign能主动暴露风险,生成全面数据集,并显著提升LVLM对复杂跨模态威胁的鲁棒性。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal perception and generation, yet their safety alignment remains a critical challenge.Existing defenses and vulnerable to multimodal jailbreaks, as visual inputs introduce new attack surfaces, reasoning chains lack safety supervision, and alignment often degrades under modality fusion.To overcome these limitation, we propose VisuoAlign, a framework for multi-modal safety alignment via prompt-guided tree search.VisuoAlign embeds safety constrains into the reasoning process through visual-textual interactive prompts, employs Monte Carlo Tree Search(MCTS) to systematically construct diverse safety-critical prompt trajectories, and introduces prompt-based scaling to ensure real-time risk detection and compliant responses.Extensive experiments demonstrate that VisuoAlign proactively exposes risks, enables comprehensive dataset generation, and significantly improves the robustness of LVLMs against complex cross-modal threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。