提出PFlowNet模型,让视觉推理更可靠且可解释。
Perceptual Flow Network for Visually Grounded Reasoning

- 分离感知与推理,自适应生成视觉路径
- 在V* Bench达90.6%、MME-RealWorld-lite达67.0%新高
- 适合需要可信视觉推理的应用场景
尽管大型视觉语言模型(LVLMs)取得成功,但通用优化目标(如标准MLE)无法约束视觉轨迹,导致语言偏见和幻觉。现有方法引入视觉专家的几何先验作为额外监督,但我们发现此类监督往往次优:过度追求几何精度,推理能力有限。为此,我们提出感知流网络(PFlowNet),摆脱对专家先验的刚性对齐,实现可解释且更有效的视觉推理。具体地,PFlowNet将感知与推理解耦,建立自条件生成过程;结合多维奖励与邻域几何塑造,通过变分强化学习,促进以推理为导向的感知行为,同时保持视觉可靠性。PFlowNet具备可证明的性能保证,在实证中表现优异,尤其在V* Bench(90.6%)和MME-RealWorld-lite(67.0%)上创下新SOTA纪录。
原文摘要 · Abstract (English)
Despite the success of Large-Vision Language Models (LVLMs), general optimization objectives (e.g., standard MLE) fail to constrain visual trajectories, leading to language bias and hallucination. To mitigate this, current methods introduce geometric priors from visual experts as additional supervision. However, we observe that such supervision is typically suboptimal: it is biased toward geometric precision and offers limited reasoning utility. To bridge this gap, we propose Perceptual Flow Network (PFlowNet), which eschews rigid alignment with the expert priors and achieves interpretable yet more effective visual reasoning. Specifically, PFlowNet decouples perception from reasoning to establish a self-conditioned generation process. Based on this, it integrates multi-dimensional rewards with vicinal geometric shaping via variational reinforcement learning, thereby facilitating reasoning-oriented perceptual behaviors while preserving visual reliability. PFlowNet delivers a provable performance guarantee and competitive empirical results, particularly setting new SOTA records on V* Bench (90.6%) and MME-RealWorld-lite (67.0%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。