arXiv:2608.14558cs.AIcs.CV2026-08

让模型从笔迹声音和手部动作推断未见文字,揭示当前模型在抽象推理上的严重不足。

The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

论文配图:The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning
图 1 · 摘自论文原文
  • 通过音视频结合推断无墨迹书写内容,挑战跨模态抽象推理能力。
  • 人类正确率超80%,顶尖模型最高仅10%左右,差距显著。
  • 双模态输入反而降低表现,暴露模型融合感知线索的深层缺陷。

当前多模态模型在静态视觉与听觉内容识别上表现优异,但在从动态生成过程推断未见信息的抽象感知推理方面仍存在重大短板。本文提出「Unwritten Benchmark」新基准,旨在评估该能力。核心任务为声-动词推断:模型需仅凭笔尖刮擦声与手部运动视频,不依赖可见墨迹,推断出三种不同书写风格的文字内容。实验结果显示,人类参与者有序字母准确率超过80%,而领先多模态模型如GPT-4o与Gemini 2.5-Pro均未能突破10%。此外,我们发现模型存在悖论性融合效应——同时提供双模态信息时性能反而下降,表明其在整合互补感知线索进行认知推理方面存在根本性缺陷。该结果凸显当前模型在跨模态因果推理及微观运动理解方面的显著局限。

原文摘要 · Abstract (English)

Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a critical and underexplored frontier. In this paper, we introduce The Unwritten Benchmark, a new challenge designed to probe this abstract perceptual and cognitive ability. We define the core task as acousto-kinematic word inference: models must decipher words, across 3 different writing styles, being written solely from the audio of pen scratches and the video of hand movements, without any visible ink trace. Our evaluation results reveal a profound gap between human and machine performance: while human participants achieve high ordered letter accuracy (over 80%), leading Multimodal Machine Learning Models, including GPT-4o and Gemini 2.5-Pro, struggle significantly, failing to surpass 10%. Furthermore, we identify a paradoxical fusion effect in the models, where providing both modalities often degrades performance rather than improving it. This finding indicates a fundamental breakdown in their ability to synthesize complementary perceptual cues for this cognitive task. These findings highlight significant limitations in both cross-modal causal reasoning and the understanding of the micro-kinematics essential for such cognitive and intuitive perceptual reasoning.

多模态抽象推理跨模态融合认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。