arXiv:2507.20188cs.CV2025-07被引 1

用语义提示增强多语言文字检测,提升复杂场景识别能力

SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection

  • 通过跨模态注意力融合文本提示与视觉特征
  • 在MLT-2019和CTW1500上分别达84.8%和90.2%的F-score
  • 适合多语言、不规则形状文字检测任务

自然场景中的文字检测仍具挑战性,尤其在多种文字和任意形状情况下,仅依赖视觉线索往往不足。现有方法未充分利用语义上下文。本文提出SAViL-Det,一种新型语义感知视觉-语言模型,通过有效整合文本提示与视觉特征来提升多语言文字检测性能。该模型采用预训练的CLIP结合渐近特征金字塔网络(AFPN)实现多尺度视觉特征融合。核心是新型语言-视觉解码器,通过跨模态注意力自适应地将细粒度语义信息从文本提示传递至视觉特征。此外,引入文本到像素的对比学习机制,显式对齐文本与对应视觉像素特征。在多个挑战性基准上的实验表明,该方法达到当前最优性能,在多语言MLT-2019数据集上F-score为84.8%,在曲线文字CTW1500数据集上达90.2%。

原文摘要 · Abstract (English)

Detecting text in natural scenes remains challenging, particularly for diverse scripts and arbitrarily shaped instances where visual cues alone are often insufficient. Existing methods do not fully leverage semantic context. This paper introduces SAViL-Det, a novel semantic-aware vision-language model that enhances multi-script text detection by effectively integrating textual prompts with visual features. SAViL-Det utilizes a pre-trained CLIP model combined with an Asymptotic Feature Pyramid Network (AFPN) for multi-scale visual feature fusion. The core of the proposed framework is a novel language-vision decoder that adaptively propagates fine-grained semantic information from text prompts to visual features via cross-modal attention. Furthermore, a text-to-pixel contrastive learning mechanism explicitly aligns textual and corresponding visual pixel features. Extensive experiments on challenging benchmarks demonstrate the effectiveness of the proposed approach, achieving state-of-the-art performance with F-scores of 84.8% on the benchmark multi-lingual MLT-2019 dataset and 90.2% on the curved-text CTW1500 dataset.

多语言检测视觉语言模型文本对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。