让AI像人一样有顺序地看图推理,提升视觉理解能力。
Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning
- 按问题重要性排序视觉区域,模拟人类逐层观察
- 在多个视觉推理任务上准确率显著提升
- 无需额外标注,适合通用多模态模型改进
当前多模态大模型将图像编码为静态视觉前缀,依赖文本推理,缺乏目标驱动和自适应视觉访问。受人类视觉感知启发——注意力从最相关信息逐步转移到次要线索——我们提出结构化序列视觉思维链(SSV-CoT)。首先,基于问题相关性生成显著性图,组织关键视觉区域,显式建模视觉重要性的空间分布;其次,按此判别顺序进行推理,形成从主到次的语义渐进过程。该方法端到端训练,仅需文本思维链与答案监督,不依赖区域级标注或外部工具。在多个视觉推理基准测试中表现提升,验证了结构化、序列化视觉认知的有效性。
原文摘要 · Abstract (English)
Current multimodal LLMs encode images as static visual prefixes and rely on text-based reasoning, lacking goal-driven and adaptive visual access. Inspired by human visual perception-where attention is selectively and sequentially shifted from the most informative regions to secondary cues-we propose Structural Sequential Visual CoT SSV-CoT. First, a question-relevant saliency map identifies and organizes key visual regions, explicitly modeling the spatial distribution of visual importance. Second, reasoning is performed following this discriminative order, inducing a curriculum-like semantic progression from primary to secondary cues. This method is trained end-to-end, using text cot and answer supervision, without relying on region-level annotations or specialized external tools. Experiments on diverse visual reasoning benchmarks show gains, validating structured and sequential visual cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。