让大模型直接看图写摘要,推理更准更自然。
End-to-End Chart Summarization via Visual Chain-of-Thought in Vision-Language Models
- 用视觉思维链引导大模型逐步看图分析,不依赖额外解析模块。
- 在Chart-Sum-QA数据集上多项指标领先,人类评估匹配度和推理正确率更高。
- 适合需要自动解读图表的科研、商业场景,尤其看重逻辑准确性的用户。
自动化图表摘要对提升数据可访问性和高效提取视觉信息至关重要。尽管视觉语言模型(VLMs)取得进展,现有方法仍存在生成摘要与图表数据不匹配、难以理解复杂图表模式等问题。本文提出端到端视觉思维链(V-CoT)方法,针对大视觉语言模型(LVLMs)优化,直接训练其从图表图像生成文本摘要,无需显式图表解析模块。通过指令微调引入视觉思维链机制,隐式引导模型在生成摘要时进行视觉推理步骤。在大规模Chart-Sum-QA数据集上的评估显示,该方法在BLEU、BLEURT、CIDEr和CS等多项自动指标上显著优于当前最优基线,并在人类评估中表现出更高的摘要匹配度和推理正确性。消融实验与详细分析进一步验证了方法的有效性与鲁棒性,确立了端到端图表摘要的新基准。
原文摘要 · Abstract (English)
Automated chart summarization is crucial for enhancing data accessibility and enabling efficient information extraction from visual data. While recent advances in visual-language models (VLMs) have demonstrated promise, existing methods often suffer from limitations in matching the generated summary to the chart data and in reasoning about complex chart patterns. This paper introduces End-to-End Visual Chain-of-Thought (V-CoT) for chart summarization, a novel approach optimized for Large Vision-Language Models (LVLMs). Our method directly trains an LVLM to process chart images and generate textual summaries in an end-to-end fashion, eliminating the need for explicit chart parsing modules. We incorporate a visual Chain-of-Thought mechanism through instruction fine-tuning, implicitly guiding the LVLM to perform visual reasoning steps during summary generation. Evaluated on the large-scale Chart-Sum-QA dataset, our V-CoT method significantly outperforms state-of-the-art baselines across a range of automatic metrics, including BLEU, BLEURT, CIDEr, and CS, and demonstrates superior matching degree and reasoning correctness in human evaluations. Ablation studies and detailed analyses further validate the effectiveness and robustness of our proposed approach, establishing a new benchmark for end-to-end chart summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。