arXiv:2508.03164cs.CVcs.AI2025-08ICCV被引 8

构建56.5万张真实图表数据集,减少模型幻觉生成。

ChartCap: Mitigating Hallucination of Dense Chart Captioning

  • 基于四阶段流程,仅依据图表可辨识数据生成描述。
  • 引入视觉一致性评分,无需参考答案评估生成质量。
  • 微调后模型准确率超人类标注,显著降低幻觉。

为提升图表描述生成的准确性与信息量,同时减少幻觉问题,现有视觉语言模型面临缺乏大规模高质量真实图表数据集的挑战。当前真实图表数据集常包含无法从图表中推断出的冗余信息,且未能充分捕捉结构元素与关键洞察。为此,本文提出 ChartCap,一个包含 56.5 万张真实世界图表图像的大型数据集,其配对的密集描述具备类型特定性,不包含额外信息,并详细呈现结构要素与核心洞察。构建 ChartCap 采用四阶段流水线,仅利用图表可辨识数据生成描述,并引入基于循环一致性的真人验证机制,在不牺牲准确性的前提下加速质量控制。此外,提出一种新指标——视觉一致性分数(Visual Consistency Score),通过比较从描述重构的图表与原始图表之间的相似度来评估质量,独立于参考文本。大量实验证明,基于 ChartCap 微调的模型持续生成更准确、更丰富的描述,幻觉显著减少,性能超越开源与闭源模型,甚至优于人工标注结果。

原文摘要 · Abstract (English)

Generating accurate, informative, and hallucination-free captions for charts remains challenging for vision language models, primarily due to the lack of large-scale, high-quality datasets of real-world charts. However, existing real-world chart datasets suffer from the inclusion of extraneous information that cannot be inferred from the chart and failure to sufficiently capture structural elements and key insights. Therefore, we introduce ChartCap, a large-scale dataset of 565K real-world chart images paired with type-specific, dense captions that exclude extraneous information and highlight both structural elements and key insights in detail. To build ChartCap, we design a four-stage pipeline that generates captions using only the discernible data from the chart and employ a cycle consistency-based human verification, which accelerates quality control without sacrificing accuracy. Additionally, we propose a novel metric, the Visual Consistency Score, which evaluates caption quality by measuring the similarity between the chart regenerated from a caption and the original chart, independent of reference captions. Extensive experiments confirms that models fine-tuned on ChartCap consistently generate more accurate and informative captions with reduced hallucinations, surpassing both open-source and proprietary models and even human-annotated captions.

图表生成幻觉抑制数据集构建多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。