用双路径框架识别误导性图表中的视觉陷阱,提升问答准确率。
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
- 分诊断视觉路径与OCR数据路径,分别检测异常结构和核对数值。
- 在两个基准上分别达到74.43%和64.55%准确率,领先基线近29%。
- 适合需要高可靠性图表分析的场景,如金融、科研报告审查。
尽管视觉语言模型(VLMs)取得进展,误导性图表仍因欺骗性视觉结构和失真数据表示构成重大挑战。我们提出ChartCynics,一种基于代理的双路径框架,通过‘质疑性’推理机制揭露视觉欺骗。不同于整体式模型,ChartCynics将感知与验证分离:诊断视觉路径通过策略性区域裁剪捕捉结构异常(如坐标轴反转),而OCR驱动的数据路径确保数值真实性。为解决跨模态冲突,引入经两阶段优化的代理摘要器:基于模拟专家的SFT用于推理提炼,基于欺骗感知的GRPO实现对抗性对齐。该流程有效惩罚视觉诱饵并强制逻辑一致性。在两个基准上的评估显示,ChartCynics分别达到74.43%和64.55%准确率,相较Qwen3-VL-8B骨干模型绝对提升约29%,超越当前最先进专有模型。结果表明,专用代理工作流可使小型开源模型具备更优鲁棒性,为可信图表解析建立新基础。
原文摘要 · Abstract (English)
Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework designed to unmask visual deception via a "skeptical" reasoning paradigm. Unlike holistic models, ChartCynics decouples perception from verification: a Diagnostic Vision Path captures structural anomalies (e.g., inverted axes) through strategic ROI cropping, while an OCR-Driven Data Path ensures numerical grounding. To resolve cross-modal conflicts, we introduce an Agentic Summarizer optimized via a two-stage protocol: Oracle-Informed SFT for reasoning distillation and Deception-Aware GRPO for adversarial alignment. This pipeline effectively penalizes visual traps and enforces logical consistency. Evaluations on two benchmarks show that ChartCynics achieves 74.43% and 64.55% accuracy, providing an absolute performance boost of ~29% over the Qwen3-VL-8B backbone, outperforming state-of-the-art proprietary models. Our results demonstrate that specialized agentic workflows can grant smaller open-source models superior robustness, establishing a new foundation for trustworthy chart interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。