让超分辨率图像既清晰又可读,特别修复模糊文字。
GLYPH-SR: Can We Achieve Both High-Quality Image Super-Resolution and High-Fidelity Text Recovery via VLM-guided Latent Diffusion Model?
- 用OCR引导的扩散模型,分阶段优化文字和画面质量。
- 在SVT等数据集上,文字识别准确率最高提升15.18个百分点。
- 适合需要精准文字识别的场景,如文档、零售分析。
图像超分辨率(SR)是监控、自动驾驶、文档分析和零售分析等视觉系统的基础,因恢复高频细节(尤其是场景文字)可保障下游感知可靠性。场景文字(如招牌、标签、店招)常携带关键信息;当字符模糊或生成错误时,即使图像整体清晰,光学字符识别(OCR)及后续决策也会失败。然而,以往的SR研究多依赖对失真度(PSNR/SSIM)或感知指标(LIPIS、MANIQA、CLIP-IQA、MUSIQ),这些指标对字符级错误不敏感。此外,关注文字SR的研究常局限于孤立字符的简化基准,忽略了复杂自然场景中的挑战,导致场景文字被当作普通纹理处理。因此,实际部署中必须同时优化文字可读性与视觉感知质量。本文提出GLYPH-SR,一种视觉-语言引导的扩散框架,联合实现双重目标。该框架采用由OCR数据引导的文本-超分融合控制网(TS-ControlNet),以及在文本与场景导向间交替的乒乓调度器。为实现针对性文字修复,我们基于合成语料训练这些组件,同时冻结主超分分支。在SVT、SCUT-CTW1500和CUTE80数据集上,于x4和x8缩放下,GLYPH-SR相较扩散/生成对抗网络基线(如SVT x8, OpenOCR),OCR F1值最高提升+15.18个百分点,同时保持竞争性的MANIQA、CLIP-IQA和MUSIQ表现。GLYPH-SR旨在同步满足高可读性与高视觉真实感,实现既‘看起来对’也‘读得懂’的超分辨率效果。
原文摘要 · Abstract (English)
Image super-resolution(SR) is fundamental to many vision system-from surveillance and autonomy to document analysis and retail analytics-because recovering high-frequency details, especially scene-text, enables reliable downstream perception. Scene-text, i.e., text embedded in natural images such as signs, product labels, and storefronts, often carries the most actionable information; when characters are blurred or hallucinated, optical character recognition(OCR) and subsequent decisions fail even if the rest of the image appears sharp. Yet previous SR research has often been tuned to distortion (PSNR/SSIM) or learned perceptual metrics (LIPIS, MANIQA, CLIP-IQA, MUSIQ) that are largely insensitive to character-level errors. Furthermore, studies that do address text SR often focus on simplified benchmarks with isolated characters, overlooking the challenges of text within complex natural scenes. As a result, scene-text is effectively treated as generic texture. For SR to be effective in practical deployments, it is therefore essential to explicitly optimize for both text legibility and perceptual quality. We present GLYPH-SR, a vision-language-guided diffusion framework that aims to achieve both objectives jointly. GLYPH-SR utilizes a Text-SR Fusion ControlNet(TS-ControlNet) guided by OCR data, and a ping-pong scheduler that alternates between text- and scene-centric guidance. To enable targeted text restoration, we train these components on a synthetic corpus while keeping the main SR branch frozen. Across SVT, SCUT-CTW1500, and CUTE80 at x4, and x8, GLYPH-SR improves OCR F1 by up to +15.18 percentage points over diffusion/GAN baseline (SVT x8, OpenOCR) while maintaining competitive MANIQA, CLIP-IQA, and MUSIQ. GLYPH-SR is designed to satisfy both objectives simultaneously-high readability and high visual realism-delivering SR that looks right and reds right.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。