arXiv:2603.22819cs.CVcs.AI2026-03中稿 · CVPR被引 1

通过细节感知与单元格对齐,提升小数据下的表格识别效果

TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment

  • 先感知后融合:用语言建模方式联合学习结构与内容
  • 在有限数据下仍达顶尖性能,无需特定数据微调
  • 新增结构引导的单元格定位模块,提升准确率与可解释性

表格广泛存在于各类文档中,表格识别(TR)是文档分析的基础任务。现有模块化流程分别建模表格结构与内容,导致整合不佳且流程复杂。端到端方法依赖大规模数据,在数据受限场景表现欠佳。为此,本文提出TDATR(Table Detail-Aware Table Recognition),通过表格细节感知学习与单元格级视觉对齐,改进端到端表格识别。TDATR采用“感知-融合”策略:首先在语言建模范式下设计多任务,联合感知表格结构与内容,自然利用多样场景文档数据增强模型鲁棒性;随后融合隐式表格细节生成结构化HTML输出,实现小样本高效建模。此外,设计结构引导的单元格定位模块,嵌入端到端框架,精准定位单元格并强化视觉-语言对齐,提升可解释性与准确性。在七个基准上达到最先进或极具竞争力的表现,无需特定数据微调。

原文摘要 · Abstract (English)

Tables are pervasive in diverse documents, making table recognition (TR) a fundamental task in document analysis. Existing modular TR pipelines separately model table structure and content, leading to suboptimal integration and complex workflows. End-to-end approaches rely heavily on large-scale TR data and struggle in data-constrained scenarios. To address these issues, we propose TDATR (Table Detail-Aware Table Recognition) improves end-to-end TR through table detail-aware learning and cell-level visual alignment. TDATR adopts a ``perceive-then-fuse'' strategy. The model first performs table detail-aware learning to jointly perceive table structure and content through multiple structure understanding and content recognition tasks designed under a language modeling paradigm. These tasks can naturally leverage document data from diverse scenarios to enhance model robustness. The model then integrates implicit table details to generate structured HTML outputs, enabling more efficient TR modeling when trained with limited data. Furthermore, we design a structure-guided cell localization module integrated into the end-to-end TR framework, which efficiently locates cell and strengthens vision-language alignment. It enhances the interpretability and accuracy of TR. We achieve state-of-the-art or highly competitive performance on seven benchmarks without dataset-specific fine-tuning.

表格识别视觉语言小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。