arXiv:2512.06759cs.CVcs.AI2025-12被引 3

评测大模型在多图多轮视觉推理中摆脱语言依赖的能力

VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors

  • 用多智能体生成框架构建高多样性图像任务
  • 涵盖20,000张图像,1,457个需逐步推理的任务
  • 适合评估真实场景下视觉链式推理能力的模型

理解多图像、多轮次场景是大型视觉语言模型(LVLMs)的关键能力,但目前仍研究不足。现有基准大多聚焦静态或横向比较(如发现视觉差异或判断适当性),且过度依赖语言线索。这类设置忽略了渐进式、上下文相关的推理以及视觉到视觉的推断挑战。为此,我们提出VisChainBench,一个大规模基准,用于严格评估LVLM在序列化、相互依赖的任务中进行多步视觉推理的能力,且语言引导极小。该基准包含1,457个任务,覆盖超过20,000张图像,涉及多个领域(如日常场景、工程故障排查),结构设计模拟真实决策过程。其独特之处在于采用多智能体生成流水线,确保高视觉多样性和可控的语言偏差。所有基准数据及构造代码已公开,可通过链接访问:https://huggingface.co/datasets/eyehole/VisChainBench

原文摘要 · Abstract (English)

Understanding multi-image, multi-turn scenarios is a critical yet underexplored capability for Large Vision-Language Models (LVLMs). Existing benchmarks predominantly focus on static or horizontal comparisons -- e.g., spotting visual differences or assessing appropriateness -- while relying heavily on language cues. Such settings overlook progressive, context-dependent reasoning and the challenge of visual-to-visual inference. To bridge this gap, we present VisChainBench, a large-scale benchmark designed to rigorously evaluate LVLMs' ability to perform multi-step visual reasoning across sequential, interdependent tasks with minimal language guidance. VisChainBench contains 1,457 tasks spanning over 20,000 images across three diverse domains (e.g., daily scenarios, engineering troubleshooting), structured to mimic real-world decision-making processes. Uniquely, the benchmark is constructed using a multi-agent generation pipeline, ensuring high visual diversity and controlled language bias. All the benchmark data and code for benchmark construction are available for viewing and download via following Link: https://huggingface.co/datasets/eyehole/VisChainBench

视觉推理多模态基准测试大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。