通过观察反馈优化推理路径,提升多步视觉推理准确性
V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning

- 引入思考-执行-观察迭代框架,动态调整推理策略
- 在8个基准上平均提升19.7%,显著优于基线模型
- 适合需要高精度多步视觉推理的应用场景
多模态大语言模型在通用感知任务中表现优异,但复杂多步视觉推理仍具挑战。现有代理式方法常忽略执行反馈,导致想象-行动-观察(IAO)偏差,影响推理稳定性和最优性。为此,我们提出V-ABS,一种基于动作-观察的束搜索框架,通过思考-执行-观察迭代实现可控推理。设计熵驱动的自适应加权算法,动态平衡策略先验与观测反馈的置信度。同时构建包含超8万样本的监督微调数据集,引导模型对正确动作路径赋予更高先验置信。在8个多样化基准上的实验证明,V-ABS性能达当前最优,相较Qwen3-VL-8B基线平均提升19.7%,且在开源与闭源模型上均保持一致增益。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have achieved remarkable success in general perception, yet complex multi-step visual reasoning remains a persistent challenge. Although recent agentic approaches incorporate tool use, they often neglect critical execution feedback. Consequently, they suffer from the imagination-action-observer (IAO) bias, a misalignment between prior imagination and observer feedback that undermines reasoning stability and optimality. To bridge this gap, we introduce V-ABS, an action-observer driven beam search framework that enables deliberate reasoning through thinker-actor-observer iterations. We also propose an entropy-based adaptive weighting algorithm to mitigate the IAO bias by dynamically balancing the confidence scores between the policy priors and the observational feedback. Moreover, we construct a large-scale supervised fine-tuning (SFT) dataset comprising over 80k samples to guide the model to assign higher prior confidence to correct action paths. Extensive experiments across eight diverse benchmarks show that V-ABS achieves state-of-the-art performance, delivering an average improvement of 19.7% on the Qwen3-VL-8B baseline and consistent gains across both open-source and proprietary models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。