用视觉锚定的思维链提升多模态推理,解决图文幻觉问题。
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
- 基于7.1万张图表的分步视觉锚定推理训练
- 在ChartQA上准确率从70.88%提升至90.04%
- 适用于复杂真实场景的跨域图文理解任务
近期大语言模型通过链式思维(CoT)和强化学习显著提升了文本推理能力,但将其扩展到视觉-语言任务仍面临挑战,主要源于纯文本CoT存在视觉幻觉和多模态融合不足的问题。本文提出Point-RFT,一种专为视觉文档理解设计的多模态推理框架,采用两阶段策略:首先利用包含7.1万条多样化视觉推理题目的数据集进行格式微调,每道题均标注与视觉元素明确关联的逐步推理过程;其次针对视觉文档理解任务实施强化微调。在ChartQA上,该方法将准确率从格式微调基线的70.88%提升至90.04%,超过仅依赖文本CoT的强化微调所达的83.92%。结果表明,视觉锚定的链式思维比纯文本链更有效。此外,Point-RFT在CharXiv、PlotQA、IconQA、TabMWP等多个跨域视觉文档推理基准上展现出优越泛化能力,凸显其在复杂现实场景中的潜力。
原文摘要 · Abstract (English)
Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language tasks remains challenging due to inherent limitations in text-only CoT, such as visual hallucinations and insufficient multimodal integration. In this paper, we introduce Point-RFT, a multimodal reasoning framework explicitly designed to leverage visually grounded CoT reasoning for visual document understanding. Our approach consists of two stages: First, we conduct format finetuning using a curated dataset of 71K diverse visual reasoning problems, each annotated with detailed, step-by-step rationales explicitly grounded to corresponding visual elements. Second, we employ reinforcement finetuning targeting visual document understanding. On ChartQA, our approach improves accuracy from 70.88% (format-finetuned baseline) to 90.04%, surpassing the 83.92% accuracy achieved by reinforcement finetuning relying solely on text-based CoT. The result shows that our grounded CoT is more effective for multimodal reasoning compared with the text-only CoT. Moreover, Point-RFT exhibits superior generalization capability across several out-of-domain visual document reasoning benchmarks, including CharXiv, PlotQA, IconQA, TabMWP, etc., and highlights its potential in complex real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。