提出PRPO与MCDR-Bench,提升图表深度研究能力
Chart Deep Research in LVLMs via Parallel Relative Policy Optimization
- 并行多维奖励优化,解决数据与奖励冲突问题
- 构建基于误差唯一性的评测基准,实现客观量化评估
- 适合需要图表深度分析的科研与决策场景
随着数据科学快速发展,图表已从简单数据展示工具演变为洞察发现与决策支持的关键手段。然而,当前图表智能在深度研究方面存在明显局限,现有方法多聚焦于视觉识别或事实问答等浅层任务,难以支撑复杂推理与高层次数据分析。这一瓶颈源于两大技术难题:训练层面,后训练技术难以应对多维奖励信号干扰与异构数据梯度冲突,导致模型多维度能力发展失衡;评估层面,现有方法局限于事实检索与基础计算,无法衡量端到端分析推理等深层能力。为此,本文提出PRPO,通过跨奖励维度的并行优化与数据类型的能力分块,有效解耦异构数据与多维奖励信号间的冲突,保障优化稳定性。同时,基于“误差唯一性原则”构建MCDR-Bench,通过可控误差注入将主观生成评估转为客观错误识别,实现对深度研究能力的可量化评估。实验验证表明,PRPO与MCDR-Bench共同建立统一框架,系统性推进图表深度研究的协同训练与客观评估。
原文摘要 · Abstract (English)
With the rapid advancement of data science, charts have evolved from simple numerical presentation tools to essential instruments for insight discovery and decision-making support. However, current chart data intelligence exhibits significant limitations in deep research capabilities, with existing methods predominantly addressing shallow tasks such as visual recognition or factual question-answering, rather than the complex reasoning and high-level data analysis that deep research requires. This limitation stems from two primary technical bottlenecks: at the training level, existing post-training techniques exhibit deficiencies in handling multi-dimensional reward signal interference and heterogeneous data gradient conflicts, preventing models from achieving balanced development across multiple capability dimensions; at the evaluation level, current methods remain limited to factual retrieval and basic computation, failing to assess end-to-end analytic reasoning and other deep research capabilities. To address the training challenge, we propose PRPO, which performs parallel optimization across reward dimensions and capability partitioning across data types, effectively disentangling conflicts between heterogeneous data and multi-dimensional reward signals while ensuring optimization stability. For the evaluation challenge, we construct MCDR-Bench based on the ``error uniqueness principle," transforming subjective generation assessment into objective error identification through controllable error injection, enabling quantifiable evaluation of deep research capabilities. Experimental validation confirms that the proposed PRPO and MCDR-Bench jointly establish a unified framework that systematically advances chart deep research through enhanced collaborative training and objective evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。