arXiv:2512.15949cs.CV2025-12中稿 · WACV 2026被引 1

构建感知观测框架,评估多模态大模型的视觉理解真实性和鲁棒性。

The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs

  • 设计多维度测试垂直任务,覆盖从局部到全局的视觉理解能力。
  • 通过像素扰动与风格化幻觉系统性测试模型在干扰下的表现稳定性。
  • 超越准确率,揭示模型是否真正理解视觉内容而非依赖文本经验。

多模态大语言模型(MLLMs)虽日益强大,但其感知能力仍缺乏有效刻画。实践中多数模型仅扩展语言部分,复用几乎相同的视觉编码器(如 Qwen2.5-VL 3B/7B/72B),引发疑问:性能提升是源于真实视觉理解,还是依赖互联网规模的文本知识?现有评估侧重任务最终准确率,忽视鲁棒性、归因可信度及可控扰动下的推理能力。本文提出「感知观测仪」(The Perceptual Observatory)框架,从三个垂直方向评估 MLLMs:(i) 简单视觉任务,如人脸匹配与图文理解;(ii) 局部到全局理解,包括图像匹配、网格定位游戏与属性定位,检验通用视觉锚定能力。每个垂直方向均使用人脸与词语的真值数据集,并通过像素级增强与扩散风格化幻觉进行系统扰动。该框架突破排行榜准确率局限,揭示模型在扰动下如何保持感知锚定与关系结构,为分析当前及未来模型优劣提供原则性基础。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have yielded increasingly powerful models, yet their perceptual capacities remain poorly characterized. In practice, most model families scale language component while reusing nearly identical vision encoders (e.g., Qwen2.5-VL 3B/7B/72B), which raises pivotal concerns about whether progress reflects genuine visual grounding or reliance on internet-scale textual world knowledge. Existing evaluation methods emphasize end-task accuracy, overlooking robustness, attribution fidelity, and reasoning under controlled perturbations. We present The Perceptual Observatory, a framework that characterizes MLLMs across verticals like: (i) simple vision tasks, such as face matching and text-in-vision comprehension capabilities; (ii) local-to-global understanding, encompassing image matching, grid pointing game, and attribute localization, which tests general visual grounding. Each vertical is instantiated with ground-truth datasets of faces and words, systematically perturbed through pixel-based augmentations and diffusion-based stylized illusions. The Perceptual Observatory moves beyond leaderboard accuracy to yield insights into how MLLMs preserve perceptual grounding and relational structure under perturbations, providing a principled foundation for analyzing strengths and weaknesses of current and future models.

多模态视觉理解鲁棒性评估感知评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。