arXiv:2603.00206cs.CVcs.AI2026-03

TACIT基准测试用程序化图像评估模型视觉推理能力。

TACIT Benchmark: A Programmatic Visual Reasoning Benchmark for Generative and Discriminative Models

  • 用程序生成10个任务,覆盖6类视觉推理场景。
  • 支持生成与判别双轨评测,验证结果完全可复现。
  • 适合研究视觉推理、模型可解释性的研究人员。

现有视觉推理基准多依赖自然语言提示,评估范围狭窄或依赖主观评分(如大模型打分)。本文提出TACIT基准,包含10个任务,覆盖空间导航、抽象模式补全、因果模拟、逻辑约束满足、图论和拓扑等6类推理领域。该基准提供双轨评估:生成赛道要求模型输出可被确定性计算机视觉管道验证的解图;判别赛道采用五选一多项选择题,干扰项仅违反一项结构约束,迫使模型关注细微视觉差异而非表层线索。版本0.1.0发布6,000个谜题(共108,000张不同分辨率的PNG图像),所有生成过程种子确定、可复现。数据集、生成代码与评估工具已开源,许可为Apache 2.0,发布于HuggingFace(DOI: 10.57967/hf/7904)。

原文摘要 · Abstract (English)

Existing visual reasoning benchmarks predominantly rely on natural language prompts, evaluate narrow reasoning modalities, or depend on subjective scoring procedures such as LLM-as-judge. We introduce the TACIT Benchmark, a programmatic visual reasoning benchmark comprising 10 tasks across 6 reasoning domains: spatial navigation, abstract pattern completion, causal simulation, logical constraint satisfaction, graph theory, and topology. The benchmark provides dual-track evaluation: a generative track in which models must produce solution images verified through deterministic computer-vision pipelines, and a discriminative track offering five-way multiple choice with structurally plausible near-miss distractors. Each distractor violates exactly one structural constraint, requiring models to reason about fine-grained visual differences rather than exploit superficial cues. Version 0.1.0 distributes 6,000 puzzles (108,000 PNG images across three resolutions) with fully deterministic seeded generation and reproducible verification. The dataset, generation code, and evaluation harness are released under the Apache 2.0 license on HuggingFace (DOI: 10.57967/hf/7904).

视觉推理程序化数据可复现评估多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。