arXiv:2502.15678cs.LG2025-02ICML被引 3

用真人认知数据微调视觉语言模型,提升特定认知能力但难泛化。

Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language Models

  • 基于人类认知判断数据,对模型进行直观物理与因果推理微调。
  • 微调后模型在对应任务上表现提升,且更接近人类行为模式。
  • 但仅限特定任务,无法跨视觉特征或认知领域实现通用泛化。

预训练的视觉语言模型在视觉认知能力上仍远不及人类。为提升模型的视觉认知并使其更贴近人类行为,我们引入了视觉刺激和人类在视觉认知任务上的判断,构建了一致评估环境,系统性地考察模型在多个认知领域的表现。我们在真实数据上对模型进行直观物理和因果推理的微调,发现其在对应领域性能显著提升,同时增强了与人类行为的一致性。然而,任务特定的微调并未带来对其他视觉特征或不同认知领域的鲁棒泛化能力。

原文摘要 · Abstract (English)

Pre-trained vision language models still fall short of human visual cognition. In an effort to improve visual cognition and align models with human behavior, we introduce visual stimuli and human judgments on visual cognition tasks, allowing us to systematically evaluate performance across cognitive domains under a consistent environment. We fine-tune models on ground truth data for intuitive physics and causal reasoning and find that this improves model performance in the respective fine-tuning domain. Furthermore, it can improve model alignment with human behavior. However, we find that task-specific fine-tuning does not contribute to robust human-like generalization to data with other visual characteristics or to tasks in other cognitive domains.

视觉认知模型微调因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。