全面评估多模态大模型信任度,发现并缓解其潜在风险
Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- 构建三维框架,覆盖真实性、鲁棒性、安全等五方面信任维度
- 测试30个模型发现:多模态训练会放大基础模型风险,且多数缓解方法有副作用
- 提出推理增强的安全对齐方法,显著提升安全性与性能平衡
尽管多模态大语言模型(MLLMs)能力显著进步,其可信性仍是重大关切。现有评估与缓解方法常聚焦单一维度,忽视多模态引入的风险。为此,我们提出MultiTrust-X,一个涵盖评估、分析与缓解的综合性基准。定义三维框架,包含真实性、鲁棒性、安全、公平与隐私五个信任维度;新增多模态风险与跨模态影响两类风险类型;并从数据、模型架构、训练与推理算法角度整合多种缓解策略。基于该分类体系,MultiTrust-X包含32项任务与28个精选数据集,支持对30个开源及专有MLLM的全面评估与深入分析,涵盖8种代表性缓解方法。实验揭示当前模型存在显著漏洞:信任度与通用能力间存在差距,且多模态训练与推理会放大基础大模型的潜在风险。受控分析还发现,现有缓解方法虽在特定维度有提升,但很少能有效改善整体可信性,且许多带来意外权衡,损害模型实用性。这些发现为未来改进提供实践启示,例如推理能力有助于更好平衡安全与性能。基于此,我们提出推理增强的安全对齐(RESA)方法,通过链式思维识别潜在风险,取得当前最优效果。
原文摘要 · Abstract (English)
The trustworthiness of Multimodal Large Language Models (MLLMs) remains an intense concern despite the significant progress in their capabilities. Existing evaluation and mitigation approaches often focus on narrow aspects and overlook risks introduced by the multimodality. To tackle these challenges, we propose MultiTrust-X, a comprehensive benchmark for evaluating, analyzing, and mitigating the trustworthiness issues of MLLMs. We define a three-dimensional framework, encompassing five trustworthiness aspects which include truthfulness, robustness, safety, fairness, and privacy; two novel risk types covering multimodal risks and cross-modal impacts; and various mitigation strategies from the perspectives of data, model architecture, training, and inference algorithms. Based on the taxonomy, MultiTrust-X includes 32 tasks and 28 curated datasets, enabling holistic evaluations over 30 open-source and proprietary MLLMs and in-depth analysis with 8 representative mitigation methods. Our extensive experiments reveal significant vulnerabilities in current models, including a gap between trustworthiness and general capabilities, as well as the amplification of potential risks in base LLMs by both multimodal training and inference. Moreover, our controlled analysis uncovers key limitations in existing mitigation strategies that, while some methods yield improvements in specific aspects, few effectively address overall trustworthiness, and many introduce unexpected trade-offs that compromise model utility. These findings also provide practical insights for future improvements, such as the benefits of reasoning to better balance safety and performance. Based on these insights, we introduce a Reasoning-Enhanced Safety Alignment (RESA) approach that equips the model with chain-of-thought reasoning ability to discover the underlying risks, achieving state-of-the-art results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。