通过聚焦候选区域提升复杂文本检测准确率
Spotlight Text Detector: Spotlight on Candidate Regions Like a Camera
- 用聚焦机制精准校准候选文本区域特征
- 在多个数据集上优于现有最先进方法
- 适合处理形状多变、密集排列的场景文本
不规则轮廓是场景文本检测中的难点。尽管基于分割的方法借助灵活的像素预测取得显著进展,但地理上接近的文本重叠仍导致难以独立检测。部分缩放方法通过预测文本内核并扩展重构文本,但内核为人工构造对象,语义信息不完整,易出现漏检或误检。此外,与一般物体不同,场景文本的几何特征(长宽比、尺度、形状)差异显著,增加了准确检测难度。为此,本文提出一种高效的聚光灯文本检测器(STD),包含聚光灯校准模块(SCM)和多变量信息提取模块(MIEM)。SCM通过映射滤波器获取候选特征并精确校准,消除部分假阳性样本;MIEM设计多种形状方案,探索文本的多重几何特征,帮助提取不同空间关系,增强模型对内核区域的识别能力。消融实验验证了SCM和MIEM的有效性。大量实验表明,STD在ICDAR2015、CTW1500、MSRA-TD500和Total-Text等多个数据集上均优于现有最先进方法。
原文摘要 · Abstract (English)
The irregular contour representation is one of the tough challenges in scene text detection. Although segmentation-based methods have achieved significant progress with the help of flexible pixel prediction, the overlap of geographically close texts hinders detecting them separately. To alleviate this problem, some shrink-based methods predict text kernels and expand them to restructure texts. However, the text kernel is an artificial object with incomplete semantic features that are prone to incorrect or missing detection. In addition, different from the general objects, the geometry features (aspect ratio, scale, and shape) of scene texts vary significantly, which makes it difficult to detect them accurately. To consider the above problems, we propose an effective spotlight text detector (STD), which consists of a spotlight calibration module (SCM) and a multivariate information extraction module (MIEM). The former concentrates efforts on the candidate kernel, like a camera focus on the target. It obtains candidate features through a mapping filter and calibrates them precisely to eliminate some false positive samples. The latter designs different shape schemes to explore multiple geometric features for scene texts. It helps extract various spatial relationships to improve the model's ability to recognize kernel regions. Ablation studies prove the effectiveness of the designed SCM and MIEM. Extensive experiments verify that our STD is superior to existing state-of-the-art methods on various datasets, including ICDAR2015, CTW1500, MSRA-TD500, and Total-Text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。