arXiv:2604.18508cs.IRcs.AI2026-04被引 1

科学文献检索中,图片化文档表示法效果差,文本结构信息更重要

Document-as-Image Representations Fall Short for Scientific Retrieval

论文配图:Document-as-Image Representations Fall Short for Scientific Retrieval
图 1 · 摘自论文原文
  • 用LaTeX源码构建新基准,精准定位文本、表格、图表等结构化内容
  • 图像表示法在长文档中表现最差,文本表示法对图文查询也更优
  • 无需特殊训练,图文混合表示可超越传统图像嵌入方法

许多近期文档嵌入模型基于文档图像表示进行训练,将渲染后的页面作为图像处理。现有科学文献检索基准(如ArXivQA和ViDoRe)也默认将文档视为页面图像,隐式偏好此类表示。本文认为该范式不适用于以文本为主、多模态并存的科学文献,因关键证据常分散于文本、表格、图表等结构化元素中。为此,我们提出新基准ArXivDoc,基于论文原始LaTeX源码构建,直接访问章节、表格、图表、公式等结构化内容,支持基于特定证据类型的可控查询设计。我们在单向量与多向量检索模型上系统比较了纯文本、图像、多模态表示。结果表明:(1) 文档图像表示始终次优,尤其在文档变长时;(2) 纯文本表示最有效,即使面对图表示例,也能通过图注和上下文捕捉信息;(3) 图文交错表示优于文档图像方法,且无需额外训练。

原文摘要 · Abstract (English)

Many recent document embedding models are trained on document-as-image representations, embedding rendered pages as images rather than the underlying source. Meanwhile, existing benchmarks for scientific document retrieval, such as ArXivQA and ViDoRe, treat documents as images of pages, implicitly favoring such representations. In this work, we argue that this paradigm is not well-suited for text-rich multimodal scientific documents, where critical evidence is distributed across structured sources, including text, tables, and figures. To study this setting, we introduce ArXivDoc, a new benchmark constructed from the underlying LaTeX sources of scientific papers. Unlike PDF or image-based representations, LaTeX provides direct access to structured elements (e.g., sections, tables, figures, equations), enabling controlled query construction grounded in specific evidence types. We systematically compare text-only, image-based, and multimodal representations across both single-vector and multi-vector retrieval models. Our results show that: (1) document-as-image representations are consistently suboptimal, especially as document length increases; (2) text-based representations are most effective, even for figure-based queries, by leveraging captions and surrounding context; and (3) interleaved text+image representations outperform document-as-image approaches without requiring specialized training.

文档检索科学文献多模态结构化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。