arXiv:2512.14944cs.CV2025-12被引 4

用可自验证的拼图任务提升视觉语言模型的推理能力

PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs

  • 设计轻量拼图环境,内置验证机制,无需外部标注
  • 引入探索度信号动态调整课程难度,缓解解空间坍塌
  • 提出推理-答案一致性指标,有效提升模型推理可信度

基于可验证奖励的强化学习后训练(RLVR)已成为激发视觉-语言模型链式思维推理的有效路径,但其在视觉领域的扩展仍受限于高成本或噪声标注及对外部验证器的依赖。基于拼图的RLVR是潜在替代方案,但现有方法常将拼图奖励视为平坦或稀疏,削弱了群体相对学习信号。现有课程策略过于受限:主要依赖奖励统计,未考虑解空间探索,易导致滚动轨迹坍塌。此外,强化学习后训练可能引发推理与答案不一致问题。为此,我们提出PuzzleCraft——一个无需监督的框架,通过一组轻量级拼图环境实现以视觉为中心的可扩展RLVR。PuzzleCraft设计了三种受经典视觉预训练任务启发的拼图:PatchFit、Rotation和Jigsaw。我们引入一种结合难度与解空间分散性的探索信号的课程机制,并用于降低坍塌提示组的权重。同时,提出新后训练评估指标‘推理-答案一致性’(RAC),衡量链式思维对答案的支持程度。实验表明,探索感知课程显著提升RAC与下游性能。在广泛视觉基准上,PuzzleCraft提升了鲁棒性与推理一致性,在Qwen2.5-VL与Qwen3-VL模型上均带来持续收益。结果表明,可扩展的拼图式RLVR需兼顾难度与解空间坍塌控制,并辅以显式的连贯性增强机制。

原文摘要 · Abstract (English)

RL post-training with verifiable rewards (RLVR) has become a practical route to eliciting chain-of-thought reasoning in vision--language models (VLMs), but scaling it in the visual domain remains challenging due to costly or noisy supervision and reliance on external verifiers. Puzzle-based RLVR is a promising alternative, yet existing approaches often treat puzzle rewards as flat or sparse, which weakens group-relative learning signal. Existing curriculum strategies are overly restrictive: they rely mainly on reward statistics and do not account for exploration in the solution space, which can lead to collapsed rollout dynamics. Further, RL post-training can induce reasoning--answer inconsistency as training progresses. To address these shortcomings, we present PuzzleCraft, a supervision-free framework that scales vision-centric RLVR using a set of lightweight puzzle environments with built-in verification. PuzzleCraft instantiates three puzzles inspired by classic visual pretext tasks: PatchFit, Rotation, and Jigsaw. We introduce a curriculum that combines difficulty with an exploration signal derived from solution-space dispersion, and use it to downweight collapsed prompt groups. In addition, we introduce a new post-training metric, Reasoning-Answer Consistency (RAC), to measure the degree that the chain-of-though supports the answer, and show our exploration-aware curriculum improves RAC and downstream performance. Across a broad suite of vision-centric benchmarks, PuzzleCraft improves robustness and reasoning consistency, yielding consistent downstream gains on both Qwen2.5-VL and Qwen3-VL backbones. Overall, our results suggest that scalable puzzle-based RLVR benefits from curricula that account for both difficulty and solution-space collapse, together with explicit consistency-enhancing schemes.

视觉推理强化学习链式思维自验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。