现有模型很少用图像,但图像其实很重要。
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
- 设计新数据集让解题必须依赖图像
- 顶尖模型无法识别细微图像差异导致错误
- 拼接多图像编码器对数学推理无帮助
近期研究聚焦于多模态数学推理,尤其关注相关数据集与评测基准的构建。然而,视觉信息在推理中的作用仍被低估。我们的研究表明,现有多模态数学模型极少利用视觉信息,模型性能在图像被修改或移除后基本不变。这归因于文本信息和答案选项的主导作用,无意中引导模型得出正确答案。为改进评估方法,我们提出HC-M3D数据集,专门设计为需依赖图像解题,并包含相似但不同的图像,其变化会改变正确答案。测试显示,主流模型未能识别这些细微差异,暴露出当前视觉感知能力的局限。此外,将多种图像编码器组合以提升通用VQA能力的做法,对数学推理性能无实质提升。该发现也挑战了增强数学推理中视觉依赖性的路径。相关基准与代码已公开于https://github.com/Yufang-Liu/visual_modality_role。
原文摘要 · Abstract (English)
Recent research has increasingly focused on multimodal mathematical reasoning, particularly emphasizing the creation of relevant datasets and benchmarks. Despite this, the role of visual information in reasoning has been underexplored. Our findings show that existing multimodal mathematical models minimally leverage visual information, and model performance remains largely unaffected by changes to or removal of images in the dataset. We attribute this to the dominance of textual information and answer options that inadvertently guide the model to correct answers. To improve evaluation methods, we introduce the HC-M3D dataset, specifically designed to require image reliance for problem-solving and to challenge models with similar, yet distinct, images that change the correct answer. In testing leading models, their failure to detect these subtle visual differences suggests limitations in current visual perception capabilities. Additionally, we observe that the common approach of improving general VQA capabilities by combining various types of image encoders does not contribute to math reasoning performance. This finding also presents a challenge to enhancing visual reliance during math reasoning. Our benchmark and code would be available at \href{https://github.com/Yufang-Liu/visual_modality_role}{https://github.com/Yufang-Liu/visual\_modality\_role}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。