构建复杂跨图表问答基准,提升多模态分析能力
ChartWalker: Benchmarking the Cross-Chart RAG Task with Hierarchical Knowledge Graphs

- 用分层知识图谱组织图表信息,保持分析结构
- 设计可控制难度的多跳推理路径生成方法
- 适合研究多模态RAG与智能分析系统的学者
跨图表检索增强生成(Cross-Chart RAG)在科学、商业和政治等领域的复杂多模态分析任务中至关重要。现有基准要么聚焦于结构化表格,要么通过简单提取关键点生成跨图表问题,导致查询与证据间存在词汇重叠,推理链逻辑不一致。为此,我们提出 ChartWalker 框架,采用针对图表设计的分层知识图谱构建方法,按粒度组织实体与关系以保留分析结构;并提出结构感知采样算法,合成语义连贯的多跳推理路径,实现对问答生成难度和粒度的显式控制。基于此框架,我们发布了 ChartWalker-Bench,一个涵盖多个领域和跨图表问题类型的综合性基准。对主流 RAG 范式的全面评估揭示显著性能差距,凸显该基准的挑战性与实用性。此外,我们还提供 ChartWalker-Agent 作为代理基线,助力分析并启发未来系统设计。
原文摘要 · Abstract (English)
Cross-Chart Retrieval-Augmented Generation (RAG) is critical for complex multi-modal analytical tasks in scientific, business, and political domains. However, existing benchmarks either focus on tables, which are well-structured and textualized, or generate cross-chart questions by simply extracting key points, which often induces lexical overlap between queries and evidence and yields logically inconsistent reasoning chains. To address this, we introduce ChartWalker, a novel framework for constructing challenging cross-chart RAG tasks. ChartWalker features a hierarchical knowledge graph construction method tailored to charts, which organizes entities and relations by granularity to preserve analytical structure. We then propose a structure-aware sampling algorithm that synthesizes semantically coherent, multi-hop reasoning paths, enabling explicit control over query difficulty and granularity for QA generation. Built with this framework, we release ChartWalker-Bench, a comprehensive benchmark spanning diverse domains and cross-chart query types. Extensive evaluations across major RAG paradigms reveal significant performance gaps, underscoring the benchmark's difficulty and utility. Furthermore, we provide ChartWalker-Agent, an agentic baseline to facilitate analysis and inspire future system design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。