用多模态大模型分析视觉感知,让AI解释人类看图逻辑。
Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis
- 基于心理学原理设计无标注分析框架,引导大模型理解视觉复杂性。
- 首次系统评估多模态大模型在解释视觉感知中的可解释性表现。
- 适合关注人机交互、认知科学与模型透明度的研究者参考。
本文推进人工智能增强推理在人机交互(HCI)、心理学与认知科学领域的研究,聚焦视觉感知这一关键任务。我们探讨多模态大语言模型(MLLMs)在此领域的适用性,借鉴心理学与认知科学中关于人类视觉感知复杂性的既有理论,作为指导原则,使MLLMs能够比较和解释视觉内容。研究旨在基准测试不同可解释性原则下MLLMs的表现。不同于近期仅以深度学习模型预测视觉复杂性指标的方法,本工作不致力于构建新的预测模型,而是提出一种新颖的无标注分析框架,用于评估MLLMs作为人机交互任务中认知助手的效用,以视觉感知为例。核心目标是为量化与评估MLLMs在提升人类推理能力及揭示人工标注感知数据集偏差方面的可解释性,奠定方法论基础。
原文摘要 · Abstract (English)
In this paper, we advance the study of AI-augmented reasoning in the context of Human-Computer Interaction (HCI), psychology and cognitive science, focusing on the critical task of visual perception. Specifically, we investigate the applicability of Multimodal Large Language Models (MLLMs) in this domain. To this end, we leverage established principles and explanations from psychology and cognitive science related to complexity in human visual perception. We use them as guiding principles for the MLLMs to compare and interprete visual content. Our study aims to benchmark MLLMs across various explainability principles relevant to visual perception. Unlike recent approaches that primarily employ advanced deep learning models to predict complexity metrics from visual content, our work does not seek to develop a mere new predictive model. Instead, we propose a novel annotation-free analytical framework to assess utility of MLLMs as cognitive assistants for HCI tasks, using visual perception as a case study. The primary goal is to pave the way for principled study in quantifying and evaluating the interpretability of MLLMs for applications in improving human reasoning capability and uncovering biases in existing perception datasets annotated by humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。