arXiv:2506.04999cs.CV2025-06ICML被引 4

针对中文场景文字检索布局复杂问题,提出新基准与更优模型。

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

  • 设计多布局中文街景文字检索基准,覆盖竖排、跨行等复杂排版。
  • 新模型CSTR-CLIP在多个数据集上提升18.82%准确率,推理更快。
  • 适合关注中文视觉理解、多布局文本匹配的研究者使用。

中文场景文字检索旨在搜索包含特定中文查询文本的图像,因真实场景中中文文本布局复杂多样,该任务极具挑战性。现有方法多沿用英文检索方案,性能不佳。本文构建了面向多样化布局的中文街景文字检索基准(DL-CSVTR),专门评估不同文本排列(如竖排、跨行、部分对齐)下的检索表现。为此,提出CSTR-CLIP模型,融合全局视觉信息与多粒度对齐训练,采用两阶段训练策略,克服以往模型忽略文本区域外视觉特征及仅依赖单粒度对齐的缺陷,有效应对复杂布局。实验表明,该模型在现有基准上相比先前最优模型提升18.82%准确率,并具备更快推理速度;在DL-CSVTR上的分析进一步验证其在多种布局下的优越性。相关数据集与代码将公开,推动中文场景文字检索研究。

原文摘要 · Abstract (English)

Chinese scene text retrieval is a practical task that aims to search for images containing visual instances of a Chinese query text. This task is extremely challenging because Chinese text often features complex and diverse layouts in real-world scenes. Current efforts tend to inherit the solution for English scene text retrieval, failing to achieve satisfactory performance. In this paper, we establish a Diversified Layout benchmark for Chinese Street View Text Retrieval (DL-CSVTR), which is specifically designed to evaluate retrieval performance across various text layouts, including vertical, cross-line, and partial alignments. To address the limitations in existing methods, we propose Chinese Scene Text Retrieval CLIP (CSTR-CLIP), a novel model that integrates global visual information with multi-granularity alignment training. CSTR-CLIP applies a two-stage training process to overcome previous limitations, such as the exclusion of visual features outside the text region and reliance on single-granularity alignment, thereby enabling the model to effectively handle diverse text layouts. Experiments on existing benchmark show that CSTR-CLIP outperforms the previous state-of-the-art model by 18.82% accuracy and also provides faster inference speed. Further analysis on DL-CSVTR confirms the superior performance of CSTR-CLIP in handling various text layouts. The dataset and code will be publicly available to facilitate research in Chinese scene text retrieval.

中文文本场景检索多布局CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。