提出感知时间扩展新范式,显著提升多模态模型视觉理解精度。
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
- 通过分解感知任务为可处理子问题,实现丰富令牌的逐步感知
- 在DisTANCE基准上高精度准确率从8.0%提升至64.7%
- 合成数据与数学推理数据结合,泛化能力增强,适合视觉推理研究者
近期基于可验证奖励的强化学习推理时扩展方法显著提升了大视觉语言模型(LVLMs)的推理能力。受此启发,类似策略被应用于多模态推理,但对视觉感知的影响尚不明确。为此,我们引入DisTANCE——一个以视觉估计为核心的基准测试。评估结果表明,当前LVLMs感知精度有限,推理时扩展仅带来微弱提升。这归因于现有模型采用快速感知范式,将视觉理解视为一次性输出,未建模底层感知过程。为此,我们提出感知时间扩展(PTS),一种鼓励丰富令牌感知并分解复杂感知问题为中间可处理子问题的新范式,使感知能与推理时扩展对齐并受益。结合强化学习技术,PTS显著提升感知准确率,在DisTANCE上高精度表现从8.0%跃升至64.7%,且在域外任务上具有良好泛化性。令人意外的是,尽管PTS数据纯为合成,但与数学推理数据结合后,仍在推理和真实世界感知基准中持续增益。进一步分析显示,PTS引入更多感知相关令牌,并增强模型对图像令牌的关注。代码与数据将公开。
原文摘要 · Abstract (English)
Recent advances in inference-time scaling, particularly those leveraging reinforcement learning with verifiable rewards, have substantially enhanced the reasoning capabilities of Large Vision-Language Models (LVLMs). Inspired by this success, similar strategies have been applied to multimodal reasoning, yet their impact on visual perception remains unclear. To investigate this gap, we introduce DisTANCE, a perception-centric benchmark for visual estimation tasks. Evaluation results show that LVLMs exhibit limited estimation precision, and inference-time scaling offers only marginal gains. We attribute this to the fast perception paradigm of current LVLMs, where visual understanding is treated as a one-shot output without modeling the underlying perceptual process. To address this, we propose Perception-Time Scaling (PTS), a novel paradigm that encourages token-rich perception and decomposes complex perception problems into intermediate tractable sub-problems, thereby enabling perception to align with and benefit from inference-time scaling. Combined with reinforcement learning techniques, PTS significantly improves perception accuracy, raising high-precision performance on DisTANCE from 8.0% to 64.7%, and generalizes well to out-of-domain tasks. Surprisingly, even though PTS data are purely synthetic, combining them with math reasoning data yields consistent gains in both reasoning and real-world perception benchmarks. Further analysis reveals that PTS introduces more perception-related tokens and increases the model's attention to image tokens. Our code and data will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。