arXiv:2410.05474cs.CVcs.MM2024-10被引 33

测试大模型在真实世界图像损坏下的表现,发现其稳定性远不如人类。

R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?

  • 构建从拍摄到接收的33种真实损坏场景,覆盖7个步骤和7类低级属性
  • 包含2970组带人工标注的损毁前后图像问答对,评估20个主流模型
  • 揭示模型在真实环境中的鲁棒性严重不足,推动向实际应用演进

大型多模态模型(LMMs)在视觉任务中表现优异,但现实世界中的各类图像损坏使其难以保持理想状态,限制了实际应用。为此,我们提出R-Bench,一个聚焦于LMMs真实世界鲁棒性的基准。具体包括:(a) 建模从用户拍摄到模型接收的完整链路,涵盖33种损坏维度,按损坏序列分为7个步骤,按低级属性分为7个组;(b) 收集损坏前后的参考/扭曲图像数据集,包含2970组经人工标注的问答对;(c) 提出绝对与相对鲁棒性综合评估方法,并对20个主流LMMs进行测试。结果表明,尽管模型能正确处理原始参考图像,但在面对损毁图像时性能不稳定,与人类视觉系统相比存在显著差距。我们希望R-Bench能推动LMMs鲁棒性提升,使其从实验模拟走向真实应用。

原文摘要 · Abstract (English)

The outstanding performance of Large Multimodal Models (LMMs) has made them widely applied in vision-related tasks. However, various corruptions in the real world mean that images will not be as ideal as in simulations, presenting significant challenges for the practical application of LMMs. To address this issue, we introduce R-Bench, a benchmark focused on the **Real-world Robustness of LMMs**. Specifically, we: (a) model the complete link from user capture to LMMs reception, comprising 33 corruption dimensions, including 7 steps according to the corruption sequence, and 7 groups based on low-level attributes; (b) collect reference/distorted image dataset before/after corruption, including 2,970 question-answer pairs with human labeling; (c) propose comprehensive evaluation for absolute/relative robustness and benchmark 20 mainstream LMMs. Results show that while LMMs can correctly handle the original reference images, their performance is not stable when faced with distorted images, and there is a significant gap in robustness compared to the human visual system. We hope that R-Bench will inspire improving the robustness of LMMs, **extending them from experimental simulations to the real-world application**. Check https://q-future.github.io/R-Bench for details.

多模态模型鲁棒性评估真实世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。