arXiv:2412.18327cs.CV2024-12

解决图文中人类标注的识别难题,提升模型对密集文本图像的理解能力。

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images

  • 提出新任务HAUR,专门针对文本密集图像中的人类标注理解。
  • 构建HAUR-5数据集,涵盖五类常见人工标注,推动该领域研究。
  • 自研OCR-Mix模型在多项指标上超越现有方法,适合图文理解场景使用。

视觉问答(VQA)任务通过图像传递关键信息以回答文本问题,是现实场景中最常见的问答形式之一。尽管当前已有多种视觉-文本模型并在部分VQA任务上表现良好,但在理解文本密集图像中的人类标注方面仍存在显著局限。为此,我们提出了人类标注理解与识别(HAUR)任务。作为该工作的组成部分,我们构建了包含五种常见类型人类标注的HAUR-5数据集,并开发训练了OCR-Mix模型。通过全面的跨模型对比实验,结果表明,OCR-Mix在该任务上优于其他现有模型。相关数据集与模型将陆续公开。

原文摘要 · Abstract (English)

Vision Question Answering (VQA) tasks use images to convey critical information to answer text-based questions, which is one of the most common forms of question answering in real-world scenarios. Numerous vision-text models exist today and have performed well on certain VQA tasks. However, these models exhibit significant limitations in understanding human annotations on text-heavy images. To address this, we propose the Human Annotation Understanding and Recognition (HAUR) task. As part of this effort, we introduce the Human Annotation Understanding and Recognition-5 (HAUR-5) dataset, which encompasses five common types of human annotations. Additionally, we developed and trained our model, OCR-Mix. Through comprehensive cross-model comparisons, our results demonstrate that OCR-Mix outperforms other models in this task. Our dataset and model will be released soon .

图文理解标注识别视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。