评测大模型在开放性联想推理上的能力,发现其远不如人类。
MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- 构建11497个开放任务,模拟人类发散与聚合联想思维。
- 用LLM做裁判精准评估自由回答和推理过程,发现模型常犯幻觉。
- 适合研究通用AI、认知智能与多模态模型局限性的学者使用。
大型视觉语言模型(LVLMs)虽取得显著进展,但在幻觉和浅层模式匹配方面仍不及人类智能。本文聚焦一项基础但未被充分探索的认知能力:联想,这是创造性思维与知识整合的核心。现有基准多限于封闭式任务,难以反映真实场景下开放性联想推理的复杂性。为此,我们提出MM-OPERA,一个包含11,497个实例的系统性基准,涵盖两个开放任务:远距离物品联想(RIA)与上下文联想(ICA),其设计遵循人类心理测量学原则。该基准要求模型以自由回答和显式推理路径回应,模拟发散与收敛式思维。我们采用定制化的LLM-as-a-Judge策略,结合过程-奖励-知情判断,精确解析推理质量。对前沿LVLM的广泛实证研究,包括任务实例敏感性分析、评判策略有效性验证及跨能力、领域、语言、文化多样性分析,全面揭示了当前模型在联想推理中的局限,为实现更类人、通用的人工智能铺平道路。数据集与代码已公开于https://github.com/MM-OPERA-Bench/MM-OPERA。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have exhibited remarkable progress. However, deficiencies remain compared to human intelligence, such as hallucination and shallow pattern matching. In this work, we aim to evaluate a fundamental yet underexplored intelligence: association, a cornerstone of human cognition for creative thinking and knowledge integration. Current benchmarks, often limited to closed-ended tasks, fail to capture the complexity of open-ended association reasoning vital for real-world applications. To address this, we present MM-OPERA, a systematic benchmark with 11,497 instances across two open-ended tasks: Remote-Item Association (RIA) and In-Context Association (ICA), aligning association intelligence evaluation with human psychometric principles. It challenges LVLMs to resemble the spirit of divergent thinking and convergent associative reasoning through free-form responses and explicit reasoning paths. We deploy tailored LLM-as-a-Judge strategies to evaluate open-ended outputs, applying process-reward-informed judgment to dissect reasoning with precision. Extensive empirical studies on state-of-the-art LVLMs, including sensitivity analysis of task instances, validity analysis of LLM-as-a-Judge strategies, and diversity analysis across abilities, domains, languages, cultures, etc., provide a comprehensive and nuanced understanding of the limitations of current LVLMs in associative reasoning, paving the way for more human-like and general-purpose AI. The dataset and code are available at https://github.com/MM-OPERA-Bench/MM-OPERA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。