评估大模型在公平、伦理等人类中心原则上的表现,发现开源与闭源模型各有优劣。
HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
- 基于真实新闻图像构建32000个专家验证的图文对,映射到多个人类中心原则。
- 闭源模型在伦理和共情上表现更好,开源模型视觉理解更准且更鲁棒。
- 发现所有模型在公平性和多语言包容性上仍有明显短板,适合关注AI伦理的研究者。
尽管近期大型多模态模型(LMMs)在视觉语言任务上取得显著进展,但其与人类中心(HC)原则如公平性、伦理、包容性、同理心和鲁棒性的一致性常被忽视。现有LMM基准大多只关注准确率。我们提出HumaniBench,一个统一框架,用于在现实、社会背景丰富的视觉情境中刻画多模态模型的HC对齐情况。该框架包含32,000个来自真实新闻图像的专家验证图像-问题对,每个对都通过明确指标映射至一个或多个HC原则。对比15个顶尖LMMs发现:闭源系统在伦理、推理和同理心方面领先,而开源模型在视觉定位和抗干扰能力上更优。所有模型在公平性和多语言包容性上仍存在持续差距。链式思维提示和测试时扩展可使多个HC维度提升8%至12%。HumaniBench支持细粒度分析传统基准未捕捉的对齐权衡。
原文摘要 · Abstract (English)
Although recent large multimodal models (LMMs) show impressive progress on vision language tasks, their alignment with human centered (HC) principles such as fairness, ethics, inclusivity, empathy, and robustness is often overlooked. Existing LMM benchmarks are largely accuracy-agnostic. We present HumaniBench, a unified framework for characterizing HC alignment across realistic, socially grounded visual contexts. It contains 32,000 expert-verified image-question pairs from real-world news imagery, each mapped to one or more HC principles through explicit metrics. Comparing 15 state of the art LMMs reveals consistent trade -offs: proprietary systems lead on ethics, reasoning, and empathy, while open-source models show superior visual grounding and resilience. All models show persistent gaps in fairness and multilingual inclusivity. Chain-of-thought prompting and test-time scaling yield 8to 12 % gains on several HC dimensions. HumaniBench enables fine-grained analysis of alignment trade-offs not captured by conventional multimodal benchmarks. https://vectorinstitute.github.io/humanibench/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。