针对街景商铺招牌识别难题,提出多阶段融合方案提升复杂环境下的文字识别准确率。
First-place Solution for Streetscape Shop Sign Recognition Competition
- 采用多模态特征融合与自监督训练,增强模型对复杂文本的表征能力。
- 结合Transformer大模型与强化学习的BoxDQN技术,显著提升定位与识别精度。
- 适用于城市导航、商业分析等实际场景,适合从事视觉识别研究者参考。
将文字识别技术应用于街景商铺招牌,在地图导航、智慧城市建设及商业区价值评估等领域具有重要应用前景。然而,街景图像中的招牌设计复杂、文字风格多样,极大增加了识别难度。本团队在近期竞赛中提出一种新颖的多阶段方法,融合多模态特征、大规模自监督训练以及基于Transformer的大模型。同时引入基于强化学习的BoxDQN技术和文本校正方法,取得优异成绩。大量实验验证了该方法的有效性,展现出在复杂城市环境中提升文字识别能力的潜力。
原文摘要 · Abstract (English)
Text recognition technology applied to street-view storefront signs is increasingly utilized across various practical domains, including map navigation, smart city planning analysis, and business value assessments in commercial districts. This technology holds significant research and commercial potential. Nevertheless, it faces numerous challenges. Street view images often contain signboards with complex designs and diverse text styles, complicating the text recognition process. A notable advancement in this field was introduced by our team in a recent competition. We developed a novel multistage approach that integrates multimodal feature fusion, extensive self-supervised training, and a Transformer-based large model. Furthermore, innovative techniques such as BoxDQN, which relies on reinforcement learning, and text rectification methods were employed, leading to impressive outcomes. Comprehensive experiments have validated the effectiveness of these methods, showcasing our potential to enhance text recognition capabilities in complex urban environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。