发现大模型认知与视觉理解常不一致,提出新方法提升两者一致性。
Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding
- 定义认知与感知冲突,量化模型在文档理解中的不一致程度。
- 即使GPT-4o也仅有75.26%的认知感知一致性,表明普遍存在问题。
- 提出多模态知识一致性微调,有效减少冲突并提升性能。
多模态大语言模型(MLLM)在文档理解任务中展现出强大能力,该领域研究迅速增长且工业需求旺盛。文档理解需模型兼具感知与认知能力,但因训练数据标注噪声类型不同,当前MLLM常出现感知与认知之间的冲突。以文档VQA任务(认知)为例,模型生成的答案可能与通过OCR识别出的视觉内容不符,说明模型难以建立“所见”与“所知”间的内在联系。此类冲突挑战了认知应与感知一致的直观假设,制约了MLLM的性能与可解释性。本文将此类冲突定义为认知与感知(C&P)知识冲突,系统评估其在文档理解中的表现。分析显示,即使是领先的GPT-4o,C&P一致性也仅达75.26%。为此,我们提出多模态知识一致性微调方法,可显著降低所有测试模型的C&P冲突,并同时提升其在认知与感知任务上的表现。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand. As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities. However, due to different types of annotation noise in training, current MLLMs often face conflicts between perception and cognition. Taking a document VQA task (cognition) as an example, an MLLM might generate answers that do not match the corresponding visual content identified by its OCR (perception). This conflict suggests that the MLLM might struggle to establish an intrinsic connection between the information it "sees" and what it "understands". Such conflicts challenge the intuitive notion that cognition is consistent with perception, hindering the performance and explainability of MLLMs. In this paper, we define the conflicts between cognition and perception as Cognition and Perception (C&P) knowledge conflicts, a form of multimodal knowledge conflict, and systematically assess them with a focus on document understanding. Our analysis reveals that even GPT-4o, a leading MLLM, achieves only 75.26% C&P consistency. To mitigate the C&P knowledge conflicts, we propose a novel method called Multimodal Knowledge Consistency Fine-tuning. Our method reduces C&P knowledge conflicts across all tested MLLMs and enhances their performance in both cognitive and perceptual tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。