HierVA通过分层代理实现多子图图表推理,提升复杂问题解答能力。
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning

- 分层架构:高层管理器规划,专职工作者执行推理与证据收集
- 在CharXiv数据集上超越强基线,多步推理准确率显著提升
- 支持视觉聚焦与上下文压缩,适合复杂图表分析任务
高级图表问答需要精确感知微小视觉元素,并在多个子图间进行多步推理。现有多模态大模型虽擅长理解单个图表,但在跨子图多步推理上表现不佳。我们提出HierVA,一种用于图表推理的分层视觉代理框架,该框架在联合图像-文本空间中迭代构建和更新工作上下文。高层管理器生成计划并维护仅包含关键信息的紧凑上下文,而专用工作者负责推理、搜集证据并返回结果。特别地,代理分别维护视觉与文本上下文,使用缩放工具限制视觉上下文范围。在CharXiv推理子集上的实验表明,相比强基线模型,其性能持续提升;消融实验证明分层结构、受限视觉上下文和压缩上下文带来互补增益。
原文摘要 · Abstract (English)
Advanced chart question answering requires both precise perception of small visual elements and multi-step reasoning across several subplots. While existing MLLMs are strong at understanding single plots, they often struggle with multi-step reasoning across multiple subplots. We propose HierVA, a hierarchical visual agent framework for chart reasoning that iteratively constructs and updates a working context in a joint image--text space. A high-level manager generates plans and maintains a compact context containing only key information, while specialized workers perform reasoning, gather evidence, and return results. In particular, the agent maintains separate visual and textual contexts, using a zoom-in tool to restrict the visual context. Experiments on the CharXiv reasoning subset demonstrate consistent improvements over strong multimodal baselines, and ablation studies verify that hierarchical architecture, scoped visual context, and distilled context contribute complementary gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。