测试大模型在真实场景中创造性使用工具的能力,发现其常因脱离视觉证据而犯错。
Advancing Creative Physical Intelligence in Large Multimodal Models

- 构建新基准MM-CreativityBench,评估模型在复杂场景中识别物体功能并创新使用的能力。
- 现有大模型虽能生成内容,但易忽略关键部件或虚构不存在属性,导致错误。
- 提出基于功能对齐的优化方法,显著提升模型选择正确对象和减少幻觉的能力。
大型多模态模型(LMMs)在感知与推理方面已取得快速进展,但其能力是否能泛化到开放环境中的视觉具身问题解决仍不明确。此类任务不仅需回答明确问题,更需发现物体在非明显但物理上可行的方式下如何被重新利用。这种创造性解决问题的能力是人类智能的核心,但在当前基准中尚未充分测试。为此,我们引入MM-CreativityBench,一个用于评估视觉丰富、物理受限环境中具身功能创造性工具使用的基准。每个实例提供包含候选实体及其部分的结构化视图,支持对模型迭代检查场景、识别相关功能并组合视觉与物理合理解决方案的细粒度交互式评估。实验表明,当前LMMs表现不佳,并非因为生成能力不足,而是缺乏持续的具身探索。模型常忽略相关实体、忽视关键部分,或虚构图像中不存在的属性。针对这一失效模式,我们提出具身功能对齐,将创造性工具使用视为偏好学习问题。通过直接偏好优化,鼓励模型优先选择基于视觉证据的功能-属性推理,而非虚构选项。同时引入来自功能知识库的监督信号,引导更广泛的对象探索与多轮规划。结果表明,模型在选择正确实体与部分方面实现持续提升,同时大幅减少幻觉与接地相关错误。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) have rapidly advanced in perception and reasoning; however, it remains unclear whether these capabilities generalize to discovering visually grounded solutions in open-ended environments, beyond pattern recognition. In such settings, intelligence requires more than answering well-posed questions: it involves identifying how elements in a scene can be repurposed in non-obvious yet physically feasible ways. This form of creative problem-solving is central to human intelligence, but remains largely untested in current benchmarks. To evaluate this ability, we introduce MM-CreativityBench, a benchmark for affordance-grounded creative tool use in visually rich, physically constrained environments. Each instance presents a scenario image with structured views of candidate entities and their parts, enabling fine-grained, interactive evaluation of how models iteratively inspect the scene, identify relevant affordances, and compose visually and physically grounded solutions. Our experiments show that current LMMs often fall short, not due to lack of generative capability, but because they do not sustain grounded exploration. Models often overlook relevant entities, under-examine critical parts, or hallucinate attributes not grounded in the image. Motivated by this failure mode, we propose affordance-grounded alignment, which casts creative tool use as a preference learning problem. Using Direct Preference Optimization, we encourage models to prefer attribute-affordance reasoning grounded in visual evidence over hallucinated alternatives. In addition, we incorporate supervision derived from an affordance knowledge base to guide broader entity exploration and multi-turn planning. Our results show consistent gains in selecting the correct entities and parts, while substantially reducing hallucination and grounding-related errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。