arXiv:2509.13151cs.CV2025-09中稿 · ICDAR 2025

TexTAR提升多语言文档文本属性识别准确率

TexTAR : Textual Attribute Recognition in Multi-domain and Multi-lingual Document Images

  • 基于上下文感知的Transformer架构,融合2D旋转位置编码
  • 在MMTAD数据集上达到最新最优性能,显著优于现有方法
  • 适合需要跨语言、跨领域文档分析的研究与应用

识别加粗、斜体、下划线和删除线等文本属性对理解文档语义、结构和视觉呈现至关重要。这些属性突出关键信息,在文档分析中具有重要意义。现有方法在噪声环境和多语言场景下面临计算效率或适应性不足的问题。为此,我们提出TexTAR,一种多任务、上下文感知的Transformer用于文本属性识别(TAR)。我们设计了新颖的数据选择管道以增强上下文感知能力,架构采用2D RoPE(旋转位置嵌入)机制,结合输入上下文实现更精准的属性预测。同时,我们构建了MMTAD数据集,涵盖法律文书、公告和教材等真实文档,具有多样化的多语言、多领域特性,并标注了文本属性。大量实验表明,TexTAR性能超越现有方法,证明上下文感知对达到顶尖文本属性识别效果的关键作用。

原文摘要 · Abstract (English)

Recognizing textual attributes such as bold, italic, underline and strikeout is essential for understanding text semantics, structure, and visual presentation. These attributes highlight key information, making them crucial for document analysis. Existing methods struggle with computational efficiency or adaptability in noisy, multilingual settings. To address this, we introduce TexTAR, a multi-task, context-aware Transformer for Textual Attribute Recognition (TAR). Our novel data selection pipeline enhances context awareness, and our architecture employs a 2D RoPE (Rotary Positional Embedding)-style mechanism to incorporate input context for more accurate attribute predictions. We also introduce MMTAD, a diverse, multilingual, multi-domain dataset annotated with text attributes across real-world documents such as legal records, notices, and textbooks. Extensive evaluations show TexTAR outperforms existing methods, demonstrating that contextual awareness contributes to state-of-the-art TAR performance.

文本属性识别多语言文档分析Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。