arXiv:2605.09900cs.AIcs.CL2026-05

用复杂结图测试视觉语言模型,发现它们看懂图却不会推理操作。

The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark

论文配图:The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark
图 1 · 摘自论文原文
  • 构建1951种纽结的85万张图库,设计4类任务验证模型理解与推理能力
  • 主流模型在多数任务上表现仅略好于随机,最佳成绩不足随机1.5倍
  • 思考模式提升有限,说明模型缺乏模拟结变换的内在机制

视觉语言模型能识别结图中的结构,却无法据此进行操作。KnotBench 包含来自1,951个素结原型(交叉数3到19)的858,318张图像,并采用以 Regina 的规范纽结标识为基准的验证协议。其14项任务涵盖四类:等价判断、变换预测、识别和跨模态对齐;图像与符号间的拆分揭示了感知与操作之间的差距。在64K输出词元预算下,对 Claude Opus 4.7 和 GPT-5(带与不带思考模式)进行评估。在56个(任务,模型)组合中,15个达到或低于随机基线,14项任务中有8项最佳得分不足随机的1.5倍。在图到符号的转录任务中,无模型生成完全正确字符串,宽松的 Regina 解码仅在100项中恢复出0至4个正确纽结。思考模式使 Claude 整体准确率提升1.65点,GPT-5 提升9.25点,但差距缩小有限。综合四类任务表明,当前模型虽具备图结构特征,却缺乏模拟变换操作的内部机制。

原文摘要 · Abstract (English)

A vision-language model can look at a knot diagram and report what it sees, yet fail to act on that structure. KnotBench pairs an 858,318-image corpus from 1,951 prime-knot prototypes (crossing numbers 3 to 19) with a protocol whose answers are checked against Regina's canonical knot signature. Its 14 tasks span four families, equivalence judgment, move prediction, identification, and cross-modal grounding; an image-versus-symbol split locates failures along the perception-operation gap. We score Claude Opus 4.7 and GPT-5, each with and without thinking, under a 64K output-token budget matched on both vendors. Across 56 (task, model) cases, 15 sit at or below a random baseline and 8 of 14 tasks have a best score under 1.5x random. On diagram-to-symbol transcription, no model produces a strictly correct string, and permissive Regina decoding recovers the knot in 0 to 4 of 100 items. Thinking-mode reasoning lifts overall accuracy by 1.65 points for Claude and 9.25 points for GPT-5, narrowing the gap only modestly. Read together, the four families suggest current vision-language models hold features of a diagram but lack apparatus to simulate moves on those features.

视觉语言模型纽结推理基准测试认知差距

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。