受人类认知启发,用三阶段流程提升复杂场景文字识别准确率
OTSNet: A Neurocognitive-Inspired Observation-Thinking-Spelling Pipeline for Scene Text Recognition
- 模仿人脑认知分三步:观察、思考、书写,统一建模文字识别
- 在Union14M-L上达83.5%准确率,9个场景刷新纪录
- 适合处理变形、遮挡严重的复杂场景文字识别任务
由于真实世界中复杂的干扰因素,场景文字识别(STR)仍具挑战性。现有框架中视觉与语言模块分离优化,导致跨模态错位引发误差传播。视觉编码器易被背景干扰吸引注意力,解码器在解析几何扭曲文字时出现空间错位,共同降低对不规则文本的识别精度。受人类视觉感知层级认知过程启发,我们提出OTSNet,一种基于观察-思考-书写三阶段神经认知范式的新型统一STR模型。其包含三大核心组件:(1) 双注意力马可龙编码器(DAME),通过差异注意力图抑制无关区域,增强判别性关注;(2) 位置感知模块(PAM)与语义量化器(SQ),通过自适应采样融合空间上下文与字形级语义抽象;(3) 多模态协同验证器(MMCV),通过视觉、语义与字符级特征的跨模态融合实现自校正。大量实验表明,OTSNet在具有挑战性的Union14M-L基准上达到83.5%平均准确率,在严重遮挡的OST数据集上达79.1%,在14个评估场景中有9个创下新纪录。
原文摘要 · Abstract (English)
Scene Text Recognition (STR) remains challenging due to real-world complexities, where decoupled visual-linguistic optimization in existing frameworks amplifies error propagation through cross-modal misalignment. Visual encoders exhibit attention bias toward background distractors, while decoders suffer from spatial misalignment when parsing geometrically deformed text-collectively degrading recognition accuracy for irregular patterns. Inspired by the hierarchical cognitive processes in human visual perception, we propose OTSNet, a novel three-stage network embodying a neurocognitive-inspired Observation-Thinking-Spelling pipeline for unified STR modeling. The architecture comprises three core components: (1) a Dual Attention Macaron Encoder (DAME) that refines visual features through differential attention maps to suppress irrelevant regions and enhance discriminative focus; (2) a Position-Aware Module (PAM) and Semantic Quantizer (SQ) that jointly integrate spatial context with glyph-level semantic abstraction via adaptive sampling; and (3) a Multi-Modal Collaborative Verifier (MMCV) that enforces self-correction through cross-modal fusion of visual, semantic, and character-level features. Extensive experiments demonstrate that OTSNet achieves state-of-the-art performance, attaining 83.5% average accuracy on the challenging Union14M-L benchmark and 79.1% on the heavily occluded OST dataset-establishing new records across 9 out of 14 evaluation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。