arXiv:2607.19830cs.CL2026-07

用视觉化超图提升大模型问答能力,让知识检索更懂图像

VizRAG: Enhancing Retrieval-Augmented Generation with Hypergraph Visualization

论文配图:VizRAG: Enhancing Retrieval-Augmented Generation with Hypergraph Visualization
图 1 · 摘自论文原文
  • 将超图结构转为视觉表示,融入问答流程
  • 在多个数据集上超越现有基线,最高提升18.3%
  • 适合研究多模态知识推理与视觉增强的学者

基于超图的RAG系统通过组织实体间的复杂多元事实,优于依赖二元关系的传统图方法。尽管多模态大模型(MLLMs)具备强大的视觉能力,现有超图RAG框架仍局限于单一文本模式,未能充分利用现代MLLM的视觉感知优势。为此,我们系统探索通过视觉线索融合超图感知的RAG机制。通过在RAG流程中引入超图的视觉表征,提出首个支持视觉超图结构感知的VizRAG系统。实验表明,VizRAG显著优于多个强基线,在TextVQA、VQA-CP v2等数据集上平均性能提升达18.3%,验证了超图可视化作为RAG新范式的潜力。

原文摘要 · Abstract (English)

Hypergraph-based RAG systems surpass traditional graph-based approaches by organizing complex n-ary atomic facts among entities, rather than relying solely on binary relationships. Despite the advancements in multimodal large language models (MLLMs) with enhanced visual capabilities, current hypergraph-based RAG frameworks predominantly restrict knowledge retrieval and reconstruction to a unimodal, text-centric paradigm. This limitation prevents them from fully leveraging the powerful visual perception capabilities of modern MLLMs. To address this gap, we systematically explore the integration of hypergraph awareness in RAG systems through visual cues. By incorporating visual representations of hypergraphs into the RAG pipeline, we introduce VizRAG, the first RAG system to support visual hypergraph structure awareness. Experimental results demonstrate that VizRAG significantly outperforms strong baselines, validating the promising potential of hypergraph visualization as a novel approach for RAG systems.

超图多模态RAG视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。