arXiv:2608.22143cs.AI2026-08

评估小模型在机械常识题上的推理能力,验证其能否像人一样理解实物关系。

Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems

论文配图:Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems
图 1 · 摘自论文原文
  • 用链式思维逐步解析机械图中接触点、支撑关系等隐含信息
  • 在BMCT测试集上,两模型正确率均超70%,体现较强常识推理能力
  • 适合关注小模型在物理常识任务中表现的研究者和教育应用开发者

定性机械问题求解(QMPS)指在无需精确计算的情况下,通过定性推理和常识知识解决机械领域问题。这类能力是人类智能的重要体现,涵盖从拧水龙头到急救、驾驶等日常及专业场景。雇主常使用贝内特机械理解测验(BMCT)评估候选人的此类能力。本文评估了两种先进多模态模型——Gemma-3和Qwen-VL——在解析机械图像时的链式思维(CoT)与最终答案表现。每幅图像蕴含真实的空间关系事实,如齿轮接触点、支撑关系与相对重量。我们通过评估思维链的连贯性、完整性与逻辑进展,以及答案与标准解的匹配度,全面检验模型的空间与常识推理能力。

原文摘要 · Abstract (English)

Qualitative mechanical problem-solving (QMPS) refers to solving qualitative problems from the mechanical domain. Qualitative problems can be solved with minimal discipline-specific information, without any robust quantitative calculation, generally by using qualitative reasoning and commonsense knowledge. QMPS is a vital aspect of human intelligence that allows us to tackle a wide range of tasks, from simple everyday ones such as turning on a tap to complex tasks in highly demanding and well-paying jobs in various fields, e.g., emergency medicine, plumbing, driving, etc. Employers often use the Bennett Mechanical Comprehension Test (BMCT) to evaluate job candidates' ability to solve such problems. In this work, we assess two state-of-the-art multimodal models, Gemma-3 and Qwen-VL, on their ability to interpret mechanical problem images by eliciting a step-by-step chain of thought (CoT) and a final answer. Each image inherently encodes ground-truth qualitative facts, such as contact points in gears, support relations, and relative weights, which we use to evaluate each model's spatial and commonsense reasoning capabilities. We assess each chain for coherence, completeness, and logical progression to assess each model's thought process, and final answers are compared to verified solutions to measure accuracy.

视觉语言模型常识推理机械理解链式思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。