arXiv:2603.13292cs.LGcs.AI2026-03被引 1

让多模态大模型在安全与有用间智能权衡。

Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs

  • 通过风险感知聚类增强视觉风险识别能力。
  • 动态加权数据增强使模型可上下文调节安全与帮助性。
  • 在多个安全评测中优于基线5%~20%,保持推理能力。

多模态大语言模型面临严峻的安全挑战,易受越狱攻击影响,或无意生成有害内容。尽管监督微调(SFT)和强化学习(RL)是主流对齐策略,但常陷入安全与实用性之间的权衡:要么过度谨慎拒绝合理请求,要么忽视跨模态交互中的潜在风险。为此,我们提出普拉格玛-VL(Pragma-VL),一种端到端对齐算法,使模型能务实权衡安全与帮助性。首先,引入冷启动SFT阶段,通过风险感知聚类增强视觉编码器的感知能力,并使用风险描述与高质量数据交织的数据集。其次,设计一个理论保证的奖励模型,采用基于查询动态加权的数据增强方法,实现上下文相关的安全-帮助性仲裁。大量实验表明,Pragma-VL在多数多模态安全基准上相较基线提升5%至20%,同时保持数学推理与知识问答等通用能力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) pose critical safety challenges, as they are susceptible not only to adversarial attacks such as jailbreaking but also to inadvertently generating harmful content for benign users. While internal safety alignment via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) is a primary mitigation strategy, current methods often face a safety-utility trade-off: they either refuse benign queries out of excessive caution or overlook latent risks in cross-modal interactions. To resolve this, we introduce Pragma-VL, an end-to-end alignment algorithm that enables MLLMs to pragmatically arbitrate between safety and helpfulness. First, we enhance visual risk perception with a novel cold-start SFT stage. This is achieved by applying risk-aware clustering to the visual encoder and using an interleaved dataset of risk descriptions and high-quality data. Second, we introduce a theoretically-guaranteed reward model that leverages synergistic learning. We train it with a novel data augmentation method that assigns dynamic weights based on the queries, enabling contextual arbitration between safety and helpfulness. Extensive experiments show that Pragma-VL effectively balances safety and helpfulness, outperforming baselines by 5% to 20% on most multimodal safety benchmarks while preserving its general capabilities in areas such as mathematics and knowledge reasoning.

多模态安全对齐大模型权衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。