arXiv:2412.16599cs.AI2024-12被引 4

测试多模态模型对方向的理解能力,发现多数表现堪忧。

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning

  • 构建了包含三类图像的方位推理基准测试CDR
  • 多数模型表现接近随机猜测,准确率不足50%
  • 结合思维链和混合数据训练可显著提升方向理解能力

方向推理对智能系统理解现实世界至关重要。现有研究多聚焦空间推理,而指南针方向推理仍被忽视。为此,我们提出针对多模态语言模型(MLMs)的指南针方向推理(CDR)基准,包含三类图像,用于测试上下左右及南北东西方向。评估显示,多数MLMs在方向推理上表现不佳,常处于随机猜测水平。直接使用CDR数据训练效果有限,因需理解真实物理规则。通过引入混合数据与思维链(CoT)微调方法,显著提升了模型在指南针方向推理上的表现,增强了其对方向关系的理解能力。

原文摘要 · Abstract (English)

Direction reasoning is essential for intelligent systems to understand the real world. While existing work focuses primarily on spatial reasoning, compass direction reasoning remains underexplored. To address this, we propose the Compass Direction Reasoning (CDR) benchmark, designed to evaluate the direction reasoning capabilities of multimodal language models (MLMs). CDR includes three types images to test spatial (up, down, left, right) and compass (north, south, east, west) directions. Our evaluation reveals that most MLMs struggle with direction reasoning, often performing at random guessing levels. Experiments show that training directly with CDR data yields limited improvements, as it requires an understanding of real-world physical rules. We explore the impact of mixdata and CoT fine-tuning methods, which significantly enhance MLM performance in compass direction reasoning by incorporating diverse data and step-by-step reasoning, improving the model's ability to understand direction relationships.

方向推理多模态基准测试思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。