arXiv:2504.09249cs.CVcs.IR2025-04被引 3

构建手写笔记理解新基准,评估模型跨模态推理能力

NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding

  • 设计多领域复杂手写笔记数据集,支持视觉语言融合与检索
  • 提出证据定位与开放域问答双任务,要求精准定位答案区域
  • 适用于研究文档理解、多模态推理的学者与工程师

学术手写笔记的理解与推理仍是文档AI中的难题,尤其在数学公式、图表和科学符号方面。现有视觉问答(VQA)基准主要针对印刷体或结构化手写文本,难以泛化到真实笔记场景。为此,我们提出NoTeS-Bank,一个用于笔记型问答中神经转录与搜索的评估基准。该基准包含多个领域的复杂笔记,要求模型处理非结构化、多模态内容。定义两个任务:(1) 基于证据的VQA,模型需检索带边界框证据的局部答案;(2) 开放域VQA,模型需先分类领域,再检索相关文档与答案。不同于依赖OCR和结构化数据的经典文档VQA数据集,NoTeS-Bank要求视觉-语言融合、检索与多模态推理。我们对先进视觉-语言模型(VLMs)和检索框架进行评测,揭示了其在结构化转录与推理上的局限性。该基准采用NDCG@5、MRR、Recall@K、IoU和ANLS等指标,建立了视觉文档理解与推理的新标准。

原文摘要 · Abstract (English)

Understanding and reasoning over academic handwritten notes remains a challenge in document AI, particularly for mathematical equations, diagrams, and scientific notations. Existing visual question answering (VQA) benchmarks focus on printed or structured handwritten text, limiting generalization to real-world note-taking. To address this, we introduce NoTeS-Bank, an evaluation benchmark for Neural Transcription and Search in note-based question answering. NoTeS-Bank comprises complex notes across multiple domains, requiring models to process unstructured and multimodal content. The benchmark defines two tasks: (1) Evidence-Based VQA, where models retrieve localized answers with bounding-box evidence, and (2) Open-Domain VQA, where models classify the domain before retrieving relevant documents and answers. Unlike classical Document VQA datasets relying on optical character recognition (OCR) and structured data, NoTeS-BANK demands vision-language fusion, retrieval, and multimodal reasoning. We benchmark state-of-the-art Vision-Language Models (VLMs) and retrieval frameworks, exposing structured transcription and reasoning limitations. NoTeS-Bank provides a rigorous evaluation with NDCG@5, MRR, Recall@K, IoU, and ANLS, establishing a new standard for visual document understanding and reasoning.

文档理解多模态手写识别视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。