用截图统一表示多模态信息,实现跨模态高效检索。
Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval
- 将文本、图表等整合为截图统一表示。
- 提出UniSE模型,支持跨模态截图检索。
- 构建MVRB基准,验证方法显著优于现有技术。
随着多模态技术的发展,以视觉形式获取信息日益受到关注。本文正式定义了一种新兴的信息检索范式——可视化信息检索(Vis-IR),即通过统一的视觉格式‘截图’联合表示文本、图像、表格和图表等多模态信息,应用于各类检索任务。我们为此做出三项关键贡献:首先,构建了大规模数据集VIRA(Vis-IR Aggregation),涵盖来自多种来源的截图,并标注为带描述和问答对的形式;其次,提出UniSE(Universal Screenshot Embeddings)系列检索模型,使截图可在任意模态间实现查询与被查询;最后,建立MVRB(Massive Visualized IR Benchmark)基准,覆盖多种任务形式与应用场景。在MVRB上的大量实验表明,现有多模态检索器存在明显不足,而UniSE展现出显著性能提升。本工作将向社区公开,为该新兴领域奠定坚实基础。
原文摘要 · Abstract (English)
With the popularity of multimodal techniques, it receives growing interests to acquire useful information in visual forms. In this work, we formally define an emerging IR paradigm called \textit{Visualized Information Retrieval}, or \textbf{Vis-IR}, where multimodal information, such as texts, images, tables and charts, is jointly represented by a unified visual format called \textbf{Screenshots}, for various retrieval applications. We further make three key contributions for Vis-IR. First, we create \textbf{VIRA} (Vis-IR Aggregation), a large-scale dataset comprising a vast collection of screenshots from diverse sources, carefully curated into captioned and question-answer formats. Second, we develop \textbf{UniSE} (Universal Screenshot Embeddings), a family of retrieval models that enable screenshots to query or be queried across arbitrary data modalities. Finally, we construct \textbf{MVRB} (Massive Visualized IR Benchmark), a comprehensive benchmark covering a variety of task forms and application scenarios. Through extensive evaluations on MVRB, we highlight the deficiency from existing multimodal retrievers and the substantial improvements made by UniSE. Our work will be shared with the community, laying a solid foundation for this emerging field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。