用文字描述代替位置标注,实现低成本高精度文本检测。
Hear the Scene: Audio-Enhanced Text Spotting
- 仅用文字转录训练,通过查询交互学习隐含位置特征。
- 在ICDAR2013上达到83.6%的检测准确率,接近全监督模型。
- 支持语音标注,适合残障人士,提升数据采集包容性。
近期场景文本检测研究多依赖精确的位置标注,但此类标注成本高、耗时长。本文提出一种新方法,仅使用文字转录进行训练,大幅降低对复杂标注的依赖。该方法采用基于查询的范式,通过文本查询与图像嵌入的交互学习隐含位置特征,并在文本识别阶段利用注意力激活图进行特征优化。为解决从零开始弱监督训练的收敛难题,引入循环课程学习策略。此外,设计了粗到细的跨注意力定位机制,提升文本实例定位精度。特别地,系统支持音频标注,显著缩短标注时间,并为视障等残障用户提供无障碍标注方式。实验表明,本方法在多个基准测试中表现优异,证明无需大量位置标注即可实现高精度文本检测。
原文摘要 · Abstract (English)
Recent advancements in scene text spotting have focused on end-to-end methodologies that heavily rely on precise location annotations, which are often costly and labor-intensive to procure. In this study, we introduce an innovative approach that leverages only transcription annotations for training text spotting models, substantially reducing the dependency on elaborate annotation processes. Our methodology employs a query-based paradigm that facilitates the learning of implicit location features through the interaction between text queries and image embeddings. These features are later refined during the text recognition phase using an attention activation map. Addressing the challenges associated with training a weakly-supervised model from scratch, we implement a circular curriculum learning strategy to enhance model convergence. Additionally, we introduce a coarse-to-fine cross-attention localization mechanism for more accurate text instance localization. Notably, our framework supports audio-based annotation, which significantly diminishes annotation time and provides an inclusive alternative for individuals with disabilities. Our approach achieves competitive performance against existing benchmarks, demonstrating that high accuracy in text spotting can be attained without extensive location annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。