构建视频文本理解新基准,测试真实场景下文字识别鲁棒性
ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement

- 构建含4639段真实视频的多质量文本感知数据集
- 发现模糊比低分辨率更影响文字识别,修复后可能改变模型推理依据
- 适合研究视频增强与文本理解融合的算法开发者
多模态大语言模型在视觉-语言理解上进展迅速,但在文本主导的视频推理中仍对输入质量敏感。现实用户视频常含运动模糊、压缩伪影、噪声和低分辨率文字,影响可靠的文字读取与下游推理。本文提出ClearText-Video(CTVid),一个大规模、面向场景文本的基准,用于研究在可控质量变化下的文本主导视频理解。CTVid包含4,639段真实世界文本丰富的第一人称视频、超过55万帧、160万条人工验证的场景文本标注,以及22万+个中英文时空问答对。每段高质量视频均提供内容匹配的降质与修复版本,支持两类任务:文本主导视频恢复和多质量视频问答。我们在18种代表性恢复方法和16种先进多模态大模型上评估了该数据集。结果表明,视觉增强不等于文字保真或推理提升:模糊比低分辨率更具破坏性;修复后的视频可能改变模型使用的文本证据;仅依赖OCR的流水线远落后于直接多模态推理。CTVid揭示了视频恢复与文本感知理解间的差距,为恢复感知、质量鲁棒的文本主导视频系统提供了严谨基础。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have recently made strong progress in visual--linguistic understanding. However, their performance on text-centric video reasoning remains highly sensitive to input quality. Real-world user-provided videos often contain motion blur, compression artifacts, noise, and low-resolution text, which impair reliable text reading and downstream reasoning. Whether MLLMs can robustly read and reason about real-world scene text under diverse quality conditions remains a fundamental open question. We introduce ClearText-Video (CTVid), a large-scale, scene-text-aware benchmark for studying text-centric video understanding under controlled quality variation. CTVid contains 4,639 real-world text-rich egocentric videos, 550K+ frames, 1.6M human-verified scene-text annotations, and 220K+ spatial/temporal question--answer pairs in Chinese and English. For each high-quality video, CTVid provides content-matched Degraded-Quality and Restored-Quality variants, supporting two task families: Text-Centric Video Restoration and Multi-Quality VideoQA. We evaluate 18 representative restoration methods and 16 state-of-the-art MLLMs on CTVid. The results show that visual enhancement does not guarantee textual fidelity or downstream reasoning gains: blur is more damaging than low resolution, restored videos can alter the textual evidence used by MLLMs, and OCR-only pipelines remain far below direct multimodal reasoning. CTVid exposes the gap between video restoration and text-grounded understanding, providing a rigorous foundation for restoration-aware, quality-robust text-centric video systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。