arXiv:2608.04111cs.CVcs.CL2026-08

测试模型能否跨不同表达形式识别抽象结构,发现普遍存在认知迁移短板。

GEB-Bench: Abstract Structures Told in Many Voices

论文配图:GEB-Bench: Abstract Structures Told in Many Voices
图 1 · 摘自论文原文
  • 以自指、怪圈等抽象结构为单位,设计多模态表达任务。
  • 模型在单一形式中识别结构能力较强,跨形式映射准确率显著下降。
  • 适合研究通用智能与跨模态理解的学者,揭示大模型抽象能力瓶颈。

模型能否识别河流三角洲与闪电的共性结构?我们提出 GEB-Bench 基准,其基本单元是抽象结构模式——自指、怪圈、莫比乌斯环等,灵感源自《哥德尔、艾舍尔、巴赫》。每个结构以四种形式呈现:自然场景、民间故事、数学定理与程序骨架;表面参数作为干扰变量不参与评分。结构、表达形式及其转换构成小规模跨模态范畴,评测任务即为其问题。评估十二个开源与专有模型后发现,抽象失败具有规律性。核心结论是:模型在单一种类中识别结构表现良好,但跨表达形式映射能力极弱;所有模型均存在此代价,仅前沿模型能接近缩小差距。两大证据支持:错误更契合设计的形式几何而非感知几何;不同厂商的前沿模型趋向相同错误答案;表面复杂度对所有识结构模型构成负担,模型容量仅提供容错空间而非免疫能力。GEB-Bench 完全可生成,并随完整流水线发布。

原文摘要 · Abstract (English)

Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstract structural motif--self-reference, a strange loop, a Mobius twist--in the spirit of Godel, Escher, Bach. Each motif is told in several voices: a natural scene whose composition is the structure, a folk story whose telling enacts it through a mechanically checkable form device, a mathematical theorem, and a programmatic skeleton; surface parameters are declared nuisance variables and never scored. Motifs, voices, and the structural changes between them form a small cross-modal category, and GEB-Bench's tasks are its questions. Evaluating twelve open and proprietary models, we find that abstraction failure is lawful. The central finding is a gap between recognition and cross-voice mapping: models identify a structure within one voice far better than they carry it across voices; every model pays this tax, and mapping strong enough to narrow it appears only at the frontier tier. Two patterns support it. Errors align more strongly with the designed formal geometry than with measured perceptual geometries, and frontier models from different vendors converge on the same wrong answers; and surface complexity taxes every model that reads structure, with capacity buying headroom rather than immunity. GEB-Bench is fully generative and released with its pipeline.

抽象推理跨模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。