提出CAVE方法,提升视觉模型在碎片化证据下的推理能力
CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning

- 通过三类信号评估中间步骤贡献,优化推理过程
- 在新构建的TRACER-Bench上性能显著提升
- 适合需要跨区域推理的复杂视觉任务研究者
视觉语言模型在通用多模态推理中表现优异,但在整合非局部视觉信息以支持语义不明确的视觉推理方面仍面临挑战。我们将其称为碎片化视觉推理问题。为此,提出基于GRPO的结构化信用分配方法CAVE,通过信念更新、证据获取和自适应聚焦控制三种互补的推理过程信号,在动作层面评估中间步骤的贡献,引导模型优化每个推理动作并学习更可靠的视觉推理策略。同时,构建TRACER-Bench数据集,涵盖四个非局部且语义易混淆的推理维度,并提供关键中间证据以监督推理路径。实验表明,CAVE在需要整合碎片化视觉证据的任务上显著提升性能,覆盖公开基准与新提出的TRACER-Bench,同时保持在通用多模态评估中的竞争力。进一步分析显示,CAVE有效增强了视觉推理能力,在长距离和深层跨区域依赖下表现出更强鲁棒性。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have achieved strong performance on general multimodal reasoning, yet remain challenged in integrating nonlocal visual information to support semantically underdetermined visual reasoning. We describe this challenge as Fragmented Visual Reasoning. To this end, we propose Credit Assignment for Visual Evidence (CAVE), a structured process-reward method based on GRPO for interleaved visual reasoning. Specifically, CAVE evaluates the contribution of intermediate steps at the action level via three complementary reasoning process signals: belief update, evidence acquisition, and adaptive focus control, thereby guiding the model to optimize each reasoning action and learn more reliable visual reasoning strategies. Meanwhile, we construct TRACER-Bench, which covers four nonlocal and semantically confusable reasoning dimensions and provides key intermediate evidence to supervise reasoning paths. Experiments demonstrate that CAVE substantially improves performance on tasks requiring fragmented visual evidence integration, covering both public benchmarks and our newly introduced TRACER-Bench, while retaining competitive performance on general multimodal evaluations. Further analyses reveal that CAVE effectively improves the visual reasoning capacity and exhibits stronger robustness under longer-range and deeper cross-region dependencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。