训练数据的表达偏差导致视觉语言模型缺乏推理能力。
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
- 从语用学视角分析大规模数据,发现关键推理信息缺失
- 空间、时间、否定、计数四类推理能力在模型中表现差
- 单纯扩大数据或模型规模无法提升推理,需专门标注
视觉语言模型(VLMs)缺乏推理能力的问题长期困扰研究界。我们认为这源于训练数据中的表达偏差:人们默认描述视觉内容时会省略推理所需隐含信息,例如“今天在比赛!”比“一张37人站在场地后的照片”更常见。我们通过语用学理论分析OpenCLIP、LLaVA-1.5和Molmo等主流VLM的数据,发现即便数据来自网络规模或合成生成,仍存在四类推理能力(空间、时间、否定、计数)代表性不足。通过一组精心设计的基准测试,我们证明:(i) VLMs在这些被抑制的推理类型上表现不佳;(ii) 与普遍认知相反,单纯扩大数据量、模型规模或多语言支持不会自动催生这些能力;但( iii ) 有针对性地收集包含隐含信息的标注则有效。研究强调,应采用更主动的数据构建方法,而非依赖规模实现推理能力涌现。
原文摘要 · Abstract (English)
The lack of reasoning capabilities in Vision-Language Models (VLMs) has remained at the forefront of research discourse. We posit that this behavior stems from a reporting bias in their training data. That is, how people communicate about visual content by default omits tacit information needed to supervise some types of reasoning; e.g., "at the game today!" is a more likely caption than "a photo of 37 people standing behind a field". We investigate the data underlying the popular VLMs OpenCLIP, LLaVA-1.5 and Molmo through the lens of theories from pragmatics, and find that reporting bias results in insufficient representation of four reasoning skills (spatial, temporal, negation, and counting), despite the corpora being of web-scale, and/or synthetically generated. With a set of curated benchmarks, we demonstrate that: (i) VLMs perform poorly on the aforementioned types of reasoning suppressed in the training data by reporting bias; (ii) contrary to popular belief, scaling data size, model size, and to multiple languages does not result in emergence of these skills by default; but, promisingly, (iii) incorporating annotations specifically collected to obtain tacit information is effective. Our findings highlight the need for more intentional training data curation methods, rather than counting on scale for emergence of reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。