测试大模型从图像中直接还原被遮蔽文字的能力
Can MLLMs "Read" What is Missing?

- 设计新基准,让模型仅凭视觉恢复被遮内容
- 2771个样本,跨语言多层级文本重建挑战大
- 适合评估模型布局理解与视觉语义融合能力
我们提出MMTR-Bench,一个用于评估多模态大模型(MLLMs)直接从视觉上下文中重构被遮蔽文本的内在能力的基准。与传统的问答任务不同,该基准不依赖显式提示,要求模型在真实场景(如文档、网页)的单页或多页输入中恢复被遮内容。这一设计将重构任务与指令遵循能力分离,可直接评估模型对版面结构的理解、视觉定位及知识整合能力。数据集包含2,771个测试样本,覆盖多种语言和不同目标长度。为此,我们提出一种分层评估协议以应对多样性。对代表性MLLM的实验表明,该基准极具挑战性,尤其在句子和段落级别的重构上表现不佳。主页地址:https://mmtr-bench-dataset.github.io/MMTR-Bench/
原文摘要 · Abstract (English)
We introduce MMTR-Bench, a benchmark designed to evaluate the intrinsic ability of Multimodal Large Language Models (MLLMs) to reconstruct masked text directly from visual context. Unlike conventional question-answering tasks, MMTR-Bench eliminates explicit prompts, requiring models to recover masked text from single- or multi-page inputs across real-world domains such as documents and webpages. This design isolates the reconstruction task from instruction-following abilities, enabling a direct assessment of a model's layout understanding, visual grounding, and knowledge integration. MMTR-Bench comprises 2,771 test samples spanning multiple languages and varying target lengths. To account for this diversity, we propose a level-aware evaluation protocol. Experiments on representative MLLMs show that the benchmark poses a significant challenge, especially for sentence- and paragraph-level reconstruction. The homepage is available at https://mmtr-bench-dataset.github.io/MMTR-Bench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。