让大模型画画来评估空间推理能力,直观发现其真实短板。
LTD-Bench: Evaluating Large Language Models by Letting Them Draw
- 用画图任务代替抽象评分,直接观察模型空间推理能力。
- 三阶段难度测试暴露顶尖模型在语言与空间映射上的严重缺陷。
- 适合关注模型真实能力、需物理世界理解的AI研究者使用。
当前大语言模型(LLM)的评估范式存在关键盲区:依赖模糊的数值指标,掩盖了模型在空间推理方面的根本缺陷,且无法直观展现模型的实际能力。这导致性能报告与实际应用能力之间产生危险脱节,尤其在需要物理世界理解的任务中。为此,我们提出LTD-Bench,一个突破性基准,通过要求模型生成点阵或可执行代码绘制图像,将评估从抽象分数转变为可直接观察的视觉输出。该方法使空间推理局限性对非专家也一目了然,弥合了统计表现与直观判断之间的鸿沟。LTD-Bench采用综合方法,包含生成任务(测试空间想象)与识别任务(评估空间感知),覆盖三个逐步提升难度的层级,系统评估语言到空间及反向映射能力。对先进模型的广泛实验揭示惊人的能力差距:即使在传统基准上表现优异的模型,在建立语言与空间概念之间的双向映射方面仍存在深层缺陷——这一根本性局限削弱了其作为真正世界模型的潜力。此外,LTD-Bench的可视化输出支持强大的诊断分析,有望用于研究模型间的相似性。
原文摘要 · Abstract (English)
Current evaluation paradigms for large language models (LLMs) represent a critical blind spot in AI research--relying on opaque numerical metrics that conceal fundamental limitations in spatial reasoning while providing no intuitive understanding of model capabilities. This deficiency creates a dangerous disconnect between reported performance and practical abilities, particularly for applications requiring physical world understanding. We introduce LTD-Bench, a breakthrough benchmark that transforms LLM evaluation from abstract scores to directly observable visual outputs by requiring models to generate drawings through dot matrices or executable code. This approach makes spatial reasoning limitations immediately apparent even to non-experts, bridging the fundamental gap between statistical performance and intuitive assessment. LTD-Bench implements a comprehensive methodology with complementary generation tasks (testing spatial imagination) and recognition tasks (assessing spatial perception) across three progressively challenging difficulty levels, methodically evaluating both directions of the critical language-spatial mapping. Our extensive experiments with state-of-the-art models expose an alarming capability gap: even LLMs achieving impressive results on traditional benchmarks demonstrate profound deficiencies in establishing bidirectional mappings between language and spatial concept--a fundamental limitation that undermines their potential as genuine world models. Furthermore, LTD-Bench's visual outputs enable powerful diagnostic analysis, offering a potential approach to investigate model similarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。