arXiv:2605.17270cs.CV2026-05中稿 · ICML

提出首个面向场景文本跟踪的结构感知框架,解决形变、模糊与细节敏感难题。

Beyond Detection: A Structure-Aware Framework for Scene Text Tracking

论文配图:Beyond Detection: A Structure-Aware Framework for Scene Text Tracking
图 1 · 摘自论文原文
  • 无检测设计,双分支协同优化结构与语义
  • 在三个基准上最高提升11.97% AUC,性能领先
  • 适合视频文本编辑、移除等动态处理任务

当前视觉目标追踪器在通用目标上表现优异,但在场景文本追踪任务中性能显著下降。尽管该问题尚未被充分研究,视频中文本追踪对动态文本操作(如分割、删除、编辑)至关重要。为此,本文首次将此任务形式化为场景文本追踪,并提出系统性解决方案。识别出三大挑战:1)视角变化导致严重几何畸变;2)不同实例间视觉歧义高;3)对细粒度结构细节高度敏感。为此,提出统一的无检测框架SymTrack,采用双分支协同设计,融合跨专家校准机制以降低语义偏差,预测令牌修正机制纠正结构失衡,并通过自适应推理引擎在运动约束下稳定预测。针对缺乏专用基准的问题,利用三个视频文本定位数据集构建高质量标注基准。大量实验表明,SymTrack在所有三个基准上均达到新最优,尤其在$\text{BOVText}_{\text{SOT}}$上比前人最佳方法提升最高达11.97% AUC。整体工作推动了高效、全面的文本追踪发展,为更通用的视频文本操作铺平道路。

原文摘要 · Abstract (English)

Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Although currently underexplored, tracking text in videos is essential for dynamic text manipulations such as segmentation, removal, and editing. To fill this gap, this paper formalizes this specific task as Scene Text Tracking and presents the first systematic work for it. We identify three primary challenges in this task: 1) severe geometric distortions from perspective shifts, 2) high visual ambiguity across different instances, and 3) high sensitivity to fine-grained structural details. To address these issues, we propose SymTrack, a unified detection-free framework with synergistic dual-branch design. It integrates a Cross-Expert Calibration mechanism to reduce semantic bias, along with a Predictive Token Rectification mechanism to correct structural imbalances, complemented by an Adaptive Inference Engine that stabilizes predictions under motion constraints. Considering the lack of dedicated benchmarks for this task, we utilize three datasets from video text spotting to construct a benchmark with high-quality annotations. Extensive experiments demonstrate that SymTrack sets the new state-of-the-art on all three benchmarks, outperforming previous best trackers by up to 11.97\% AUC on $ \text{BOVText}_{\text{SOT}} $. Overall, our work promotes efficient and thorough text tracking, paving the way toward more generalized video text manipulation.

文本追踪结构感知视频理解双分支

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。