评测大模型在可靠性和价值观上的表现,发现其存在隐私与伦理漏洞。
REVAL: A Comprehension Evaluation on Reliability and Values of Large Vision-Language Models

- 构建14.4万样本的图文问答基准,分可靠性与价值观两大维度。
- 26个模型测试显示,抗攻击能力弱、隐私保护差、道德推理不足。
- 适合关注AI安全、伦理与可信度的研究者使用。
大型视觉语言模型(LVLMs)的快速发展凸显了全面评估框架的必要性。现有基准多聚焦于感知能力、认知能力或对抗攻击安全性,但缺乏覆盖广度与深度。为此,我们提出REVAL,一个评估LVLMs可靠性和价值观的综合性基准。该基准包含超过14.4万张图像-文本配对的VQA样本,分为两部分:可靠性,评估真实性(如感知准确性和幻觉倾向)与鲁棒性(如对对抗攻击、字体攻击和图像退化的抵抗能力);价值观,评估伦理问题(如偏见与道德理解)、安全问题(如毒性与越狱漏洞)及隐私问题(如隐私意识与数据泄露)。我们评估了26个模型,涵盖主流开源与闭源模型(如GPT-4o、Gemini-1.5-Pro)。结果表明,当前模型在感知任务和毒性规避上表现良好,但在对抗场景、隐私保护与伦理推理方面存在显著弱点。这些发现揭示了未来改进的关键方向,推动更安全、可靠、符合伦理的LVLMs发展。REVAL为研究者系统评估与比较模型提供了坚实框架。
原文摘要 · Abstract (English)
The rapid evolution of Large Vision-Language Models (LVLMs) has highlighted the necessity for comprehensive evaluation frameworks that assess these models across diverse dimensions. While existing benchmarks focus on specific aspects such as perceptual abilities, cognitive capabilities, and safety against adversarial attacks, they often lack the breadth and depth required to provide a holistic understanding of LVLMs' strengths and limitations. To address this gap, we introduce REVAL, a comprehensive benchmark designed to evaluate the \textbf{RE}liability and \textbf{VAL}ue of LVLMs. REVAL encompasses over 144K image-text Visual Question Answering (VQA) samples, structured into two primary sections: Reliability, which assesses truthfulness (\eg, perceptual accuracy and hallucination tendencies) and robustness (\eg, resilience to adversarial attacks, typographic attacks, and image corruption), and Values, which evaluates ethical concerns (\eg, bias and moral understanding), safety issues (\eg, toxicity and jailbreak vulnerabilities), and privacy problems (\eg, privacy awareness and privacy leakage). We evaluate 26 models, including mainstream open-source LVLMs and prominent closed-source models like GPT-4o and Gemini-1.5-Pro. Our findings reveal that while current LVLMs excel in perceptual tasks and toxicity avoidance, they exhibit significant vulnerabilities in adversarial scenarios, privacy preservation, and ethical reasoning. These insights underscore critical areas for future improvements, guiding the development of more secure, reliable, and ethically aligned LVLMs. REVAL provides a robust framework for researchers to systematically assess and compare LVLMs, fostering advancements in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。