让图像隐空间推理更懂语义,提升复杂问题回答能力
Semantic-Enriched Latent Visual Reasoning

- 分两阶段训练:先用细粒度属性标注增强区域隐表示
- 在40万级区域标注数据上,问答准确率显著优于基线
- 适合需要精准语义理解的视觉推理任务研究者
多模态隐空间推理旨在通过直接在紧凑的隐空间中进行视觉推理来替代显式思考。然而,现有方法主要依赖视觉监督,生成的隐表示语义丰富度不足,难以支持多样化的区域级推理任务。本文提出语义增强的隐空间视觉推理(SLVR),一种两阶段学习框架,通过属性级视觉语义增强隐表示,并将其对齐到多种推理目标。第一阶段中,SLVR在细粒度属性监督下学习语义丰富的区域中心隐表示;第二阶段设计多查询组相对策略优化(M-GRPO),实现同一区域下多个查询间的隐表示对齐。为此构建了包含约40万条区域级属性标注和80万个多查询问答样本的SLV-Set数据集,并引入SV-QA基准,用于评估在语义变化下的隐空间推理性能。实验表明,相比现有基线,SLVR显著提升了隐空间视觉推理的鲁棒性和语义一致性。
原文摘要 · Abstract (English)
Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches largely rely on visual supervision and produce latent representations that lack sufficient semantic richness, limiting their ability to support diverse region-level reasoning tasks. In this work, we introduce Semantic-Enriched Latent Visual Reasoning (SLVR), a two-stage learning framework that enriches latent representations with attribute-level visual semantics and aligns them with diverse reasoning objectives. In the first stage, SLVR learns semantically enriched region-centric latents under fine-grained attribute supervision. In the second stage, we design Multi-query Group Relative Policy Optimization (M-GRPO) to align latent representations across multiple queries grounded in the same region. To support this framework, we construct SLV-Set, comprising approximately 400K region-level attribute annotations and 800K multi-query question answering samples, and introduce SV-QA, a benchmark that evaluates latent reasoning under semantic variation. Experiments demonstrate that SLVR improves the robustness and semantic consistency of latent visual reasoning compared to existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。