提出新基准与代码生成法,提升图表指代表达定位精度与多目标支持能力
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring

- 基于绘图代码生成像素级掩码,实现细粒度图表元素定位
- 支持多目标引用、多样语言线索和多种图表类型,覆盖更真实场景
- 在多个图表基准上表现优于基线,适合视觉语言模型评估与优化
指代表达定位是视觉语言模型空间理解与推理的核心任务,但现有研究主要聚焦自然图像。现有图表指代表达定位基准存在诸多局限:(1)多采用边界框,限制对细粒度图表元素的精确定位;(2)普遍假设单目标或双目标引用,难以处理多实例引用;(3)语言表达过度依赖文本或数据排序线索;(4)仅涵盖有限图表类型。为此,我们构建了一个系统性支持多定位形式、多目标引用、多样化定位线索及多种图表类型的基准。在代表性多模态大模型上的测试显示显著性能差距。我们进一步提出基于代码的合成流水线,利用绘图程序与渲染图表元素之间的内在对齐关系,生成跨图表元素类型与粒度的像素级实例掩码。使用合成掩码训练实例分割模型,并集成至通用多模态定位框架中。所提系统在本基准上持续优于基线,并在源自ChartQA的真实图表定位任务中表现出良好泛化能力。
原文摘要 · Abstract (English)
Referring expression grounding is a core problem in visual grounding and is widely used as a diagnostic of spatial grounding and reasoning in vision and language models, yet most prior work focuses on natural images. In contrast, existing chart referring expression grounding-related benchmarks remain limited: (1) they largely adopt bounding boxes, constraining localization precision for fine chart elements (2) they mostly assume a single and two referred target instances, failing to handle multi-instance target references; (3) the language expressions over-rely on textual cues or data-rank clues (4) they cover only a narrow range of chart types. To address these issues, we introduce a chart referring expression grounding benchmark that systematically supports multiple localization forms, multiple referred targets, diverse grounding cues and diverse chart types. Results across representative multimodal large models reveal a significant performance gap. We further introduce a code-driven synthesis pipeline that exploits the inherent alignment between plotting programs and rendered chart primitives to derive pixel accurate instance masks across chart element types and granularities. We train an instance segmentation model with the synthesized masks and integrate it into a general-purpose multimodal grounding framework. The resulting system consistently outperforms baselines on our benchmark and generalizes well to a ChartQA-derived real-chart grounding benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。