多给信息未必更好,精准才是关键。
When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
- 测试三种视觉语言模型在空间推理中的信息注入效果。
- 单一精确空间线索优于多个上下文叠加,无关常识会拖累表现。
- 思维链提示仅在空间定位准确时才有效,适合优化推理系统设计者。
视觉空间推理(VSR)对现代多模态模型仍具挑战性,尽管多模态架构不断进步。常见策略是在推理时注入额外信息,如显式空间线索、外部常识知识或思维链(CoT)推理指令。然而,尚不明确何时这些信息能真正提升推理能力,何时反而引入噪声。本文针对三种代表性视觉语言模型(VLMs)和两个公开基准,开展假设驱动的分析,考察:(i) 空间上下文的类型与数量,(ii) 注入常识知识的数量与相关性,(iii) 空间定位与CoT提示的交互作用。结果揭示一致模式:更多信息不等于更好表现。精准的单个空间线索优于多上下文聚合,过度或弱相关常识知识会降低性能,而CoT提示仅在空间定位足够精确时才提升准确率。研究强调选择性、任务适配的信息注入的重要性,为构建可靠的多模态推理流程提供实践指导。
原文摘要 · Abstract (English)
Visual spatial reasoning (VSR) remains challenging for modern vision-language models (VLMs), despite advances in multimodal architectures. A common strategy is to inject additional information at inference time, such as explicit spatial cues, external commonsense knowledge, or chain-of-thought (CoT) reasoning instructions. However, it remains unclear when such information genuinely improves reasoning and when it introduces noise. In this paper, we conduct a hypothesis-driven analysis of information injection for VSR across three representative VLMs and two public benchmarks. We examine (i) the type and number of spatial contexts, (ii) the amount and relevance of injected commonsense knowledge, and (iii) the interaction between spatial grounding and CoT prompting. Our results reveal a consistent pattern: more information does not necessarily yield better reasoning. Targeted single spatial cues outperform multi-context aggregation, excessive or weakly relevant commonsense knowledge degrades performance, and CoT prompting improves accuracy only when spatial grounding is sufficiently precise. These findings highlight the importance of selective, task-aligned information injection and provide practical guidance for designing reliable multimodal reasoning pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。