用统计方法分析多模态模型视觉错误,找出深度感知和结构补全的短板。
Visual Error Patterns in Multi-Modal AI: A Statistical Approach
- 通过参数与非参数模型分析几何图像误判因素
- 梯度提升模型在交叉验证中达AUC=0.85最优性能
- 揭示深度感知与不完整结构重建是主要错误根源
多模态大语言模型(如GPT-4o)在整合文本与视觉信息方面表现优异,但在处理模糊或不完整的视觉刺激时存在系统性挑战。本研究采用统计建模方法,基于包含三维、旋转及缺失面/边等特征的几何刺激数据集,运用参数法、非参数法和集成学习技术预测分类错误。非线性梯度提升模型在交叉验证中表现最佳,AUC达到0.85。特征重要性分析表明,深度感知困难与不完整结构重构问题是导致误分类的关键因素。研究证明统计方法可有效揭示多模态大模型的局限性,并为通过引入上下文推理机制改进模型架构提供可操作洞察。
原文摘要 · Abstract (English)
Multi-modal large language models (MLLMs), such as GPT-4o, excel at integrating text and visual data but face systematic challenges when interpreting ambiguous or incomplete visual stimuli. This study leverages statistical modeling to analyze the factors driving these errors, using a dataset of geometric stimuli characterized by features like 3D, rotation, and missing face/side. We applied parametric methods, non-parametric methods, and ensemble techniques to predict classification errors, with the non-linear gradient boosting model achieving the highest performance (AUC=0.85) during cross-validation. Feature importance analysis highlighted difficulties in depth perception and reconstructing incomplete structures as key contributors to misclassification. These findings demonstrate the effectiveness of statistical approaches for uncovering limitations in MLLMs and offer actionable insights for enhancing model architectures by integrating contextual reasoning mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。