让大模型回答图表时,能精准定位依据的视觉元素。
ChartLens: Fine-grained Visual Attribution in Charts
- 用分割技术识别图表对象,结合提示工程实现细粒度归因。
- 在多个领域图表上,归因准确率提升26%至66%。
- 适合关注模型可解释性与图表理解可靠性的研究者。
多模态大模型在图表理解任务中表现日益出色,但常出现幻觉,生成内容与视觉数据矛盾。为此,我们提出面向图表的后处理视觉归因方法,用于识别支持特定回答的细粒度图表元素。本文提出ChartLens,一种基于分割的图表归因算法,利用标记集提示配合多模态大模型实现细粒度归因。同时,构建了ChartVA-Eval基准,涵盖金融、政策、经济等领域的合成与真实图表,配有细粒度归因标注。评估表明,ChartLens在细粒度归因上性能提升26%-66%。
原文摘要 · Abstract (English)
The growing capabilities of multimodal large language models (MLLMs) have advanced tasks like chart understanding. However, these models often suffer from hallucinations, where generated text sequences conflict with the provided visual data. To address this, we introduce Post-Hoc Visual Attribution for Charts, which identifies fine-grained chart elements that validate a given chart-associated response. We propose ChartLens, a novel chart attribution algorithm that uses segmentation-based techniques to identify chart objects and employs set-of-marks prompting with MLLMs for fine-grained visual attribution. Additionally, we present ChartVA-Eval, a benchmark with synthetic and real-world charts from diverse domains like finance, policy, and economics, featuring fine-grained attribution annotations. Our evaluations show that ChartLens improves fine-grained attributions by 26-66%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。